The Story Behind Cerebras’ $63 Billion IPO with Founder and CEO Andrew Feldman
Summary
Cerebras’s commercial inflection arrived when models became useful enough in 2025 for inference speed to become a daily-work constraint. Feldman claims its AI computers run inference “15, 18, 20x faster than GPUs” across model sizes and origins. His categorical demand thesis: “How big is the market for slow inference? It’s zero.”
That performance rests on a long-running wager that radical gains would require an architecture unlike the GPU, not a minor modification. Cerebras built a 46,000-square-millimeter wafer-scale chip—“the size of a dinner plate”—and spent about $8 million monthly during a 2017-19 stretch when it could not make the design work. It finally yielded the chip in summer 2019.
G42’s $1 billion order was the bridge from niche supercomputing customers to hyperscale readiness. It let Cerebras transform its supply chain, deploy and battle-test large clusters, and do training and inference at a scale no internal QA lab could reproduce. Feldman contrasts roughly a dozen first-generation sales and 300 second-generation sales with “tens of thousands” expected for the third.
The immediate public-market thesis is delivery against an OpenAI deal Feldman says is north of $20 billion. OpenAI signed after trials showed Cerebras substantially outperforming alternatives; AWS subsequently agreed to deploy its systems in AWS data centers. Cerebras is trying to increase manufacturing 10x this year.
The IPO was framed less as an exit than as slightly cheaper capital, audited legitimacy and “corporate adulthood.” The hosts introduced Cerebras at roughly a $60-$63 billion market capitalization with 800-850 employees. Feldman’s differentiation claim is that Cerebras would be the first and only, for a period, AI pure play with 100% of revenue tied to this market—“no gaming, there’s no graphics, there’s no PC.”
Cerebras’s own coding adoption shows both AI’s operating leverage and its uneven distribution. Over the past eight months, per-engineer token spending went from below $1,000 to roughly $25,000-$30,000; a small cohort governs eight or 10 agents 24/7 and has moved from 10x to 100x productivity. Feldman includes himself among the rest still “limping along.”
Feldman’s larger bet is that low latency will create AI-native businesses, not merely make today’s software faster. His analogy is Netflix: faster internet did not incrementally improve DVD delivery; it helped turn Netflix into a movie studio. Likewise, once companies fundamentally reorganize work around AI, new business models and fundamental jumps in productivity should emerge.
Deep dive
Not yet available upstream; scheduled sync will retry.