Pioneers Insight Method Research Author
Back to Pioneers
Andrew Feldman
Founders 9 Curated Dialogues

Andrew Feldman

Cerebras · Co-Founder & CEO

Frontier Insights

Frontier Thesis: AI inference is bound by memory bandwidth, not raw compute. By leveraging wafer-scale on-chip SRAM, Cerebras bypasses GPU data-movement bottlenecks (HBM/CoWoS packaging), drastically slashing token costs and targeting Nvidia’s vulnerable pricing power.

Strategic Moves: Exploiting hardware scarcity to capture sovereign Gulf capital ($1B+ raised; G42 alliance) and validating extreme workloads while bypassing immediate IPO scrutiny and Nvidia’s balance-sheet lock-in.

Risks & Warnings: Severe revenue concentration (G42 alone accounts for ~87%), geopolitical exposure from export controls, enterprise adoption friction via legal/security gatekeepers, and macro demand volatility as Nvidia aggressively defends its moat.

Key Views & Dialogues

Open Source Wins, AGI Is Here, and Scorsese’s AI Toolkit with CEOs of Cerebras & Black Forest Labs

  • 🗓️ Date2026-07-10 | 🎙️ Show:All-In

Cerebras reports a $25 billion backlog as faster inference turns token capacity and data-center execution into scarce product advantages, while open models take routine and sovereign workloads. Feldman expects its architecture to improve well over 2x in 18 months; staged releases, red-teaming and recursive model improvement remain catalysts—and risks—as multimodal AI reaches film and robotics.

View Dialogue Notes & Key Takeaways
  • Cerebras says AI infrastructure is serving booked demand, not betting that customers will eventually appear. Andrew Feldman disclosed a $25 billion backlog and described customers as “trying to capture yesterday’s demand,” with football-field-sized facilities drawing more power than midsize cities. For investors, the operational challenge is keeping customers while demand outstrips the ability to build and fill data centers.

  • Reasoning models make inference speed and token availability part of product capability. Rombach says models such as “Fable” and “5/6” increasingly infer intent without a “prompt whisperer,” while his Hermes-agent experiment used a Bittensor project running Z.ai’s GLM-5.2 and debated its own research strategy before acting. Long-running jobs can already produce “amazing things”; Rombach’s hypothetical was that 15-times-faster Cerebras inference could compress “weeks or months’ worth of thinking” into a day.

  • Cerebras expects its young architecture to improve substantially faster than conventional Moore’s law in the next 18 months. Feldman’s view is “way over 2x,” because a new architecture still offers workload-specific optimization unavailable to a 20-year-old GPU design reliant on smaller fabrication geometries. Hyperscalers’ custom chips do not invalidate merchant silicon demand: their primary objective is avoiding total dependence on any supplier.

  • Open models are becoming the cost, control, and sovereignty layer beneath frontier AI. Feldman’s analogy is that users should not “take your Ferrari to the grocery store”: frontier models handle the hardest problems, while open models absorb routine G&A and the “cutting and pasting economy.” Regulated enterprises also want domestic, on-prem deployment; Cerebras therefore runs gpt-oss-120B, GLM, Kimmy, the “Quincy” set of models, proprietary OpenAI models, and customer-built models.

  • Staged release and red-teaming become defensible when a model can expose serious software flaws in hours. Rombach relayed that Palo Alto Networks found previously unknown bugs and paused other work for six weeks of patching; Feldman said giving government defenders “2 or 3 weeks to patch any obvious holes” is reasonable. He simultaneously warned against reflexive regulation and treated a future massive data breach as inevitable: “Something will happen.”

  • By every AGI definition used 20 years ago, Feldman believes the threshold has already been crossed; the open question is where recursive improvement stops. Repeatedly asking models to learn, retry, and broaden their search yields “not a little bit better answers, but vastly better answers,” compressing thousands of human-style generations into machine-speed iteration. His ledger includes real labor dislocation, but also a chance that children and their loved ones do not die of cancer and that every child receives an adaptive tutor.

  • Black Forest Labs sees generative media as the foundation of multimodal world models, not merely an image-and-video tool business. Robin Rombach traced FLUX back to latent diffusion, then toward models pretrained on images, video, and audio and combined with action prediction—the same model could help make a movie or become “a brain on a robot.” Near-term value remains human-directed: Martin Scorsese used the models to externalize a scene in his head, while IP owners can combine controlled generation, customized models, and potentially licensed fan creativity.

  • 🔗 Original source & video: Open Source Wins, AGI Is Here, and Scorsese’s AI Toolkit with CEOs of Cerebras & Black Forest Labs

Listen to full conversation →


The IPO Comeback: Why Tech Giants Are Finally Going Public | All-In Liquidity IPO Panel

  • 🗓️ Date2026-06-06 | 🎙️ Show:All-In

The IPO pendulum is shifting toward companies listing at $1 billion to $5 billion, with Planet creating roughly 90% of its post-SPAC value in years three and four. Cerebras’s $18.50 IPO followed a difficult 9.5 years, while its chip targets 15–18-times faster OpenAI workloads; orbital data centers offer a potential two-to-three-year cost catalyst, but distributed clustering remains unresolved.

View Dialogue Notes & Key Takeaways
  • The panel’s capital-markets call is that the IPO pendulum is swinging back toward companies listing at $1 billion, $3 billion, or $5 billion instead of “stay private forever.” Planet went public at $2 billion via SPAC in 2021, with roughly 90% of its subsequent value created in years three and four. Earlier listings can transfer more upside—and more operating scrutiny—to public investors.

  • Cerebras demonstrates both IPO friction and its payoff: Andrew Feldman says “not a damn thing changes in the important parts of your business,” while Brad Gerstner described 9.5 difficult years followed by 12 easy months. The IPO priced at $18.50 after the range was taken up twice, Brad said he thought the stock opened at $32, and it was later at $23, implying a $5–6 billion market cap.

  • Planet’s thesis is that daily, global satellite imagery becomes substantially more valuable when AI turns it into answers rather than another specialized dataset. Its roughly 200-satellite fleet images the entire Earth every day, creating a historical time series for agriculture, energy, disaster response, and security; Marshall estimates a $75 billion-$100 billion Earth-observation opportunity, with AI on top.

  • Marshall expects orbital data centers to become cheaper than terrestrial facilities once launch costs fall from just over $1,000 per kilogram to roughly $200-$300, potentially within two to three years. Constant sunlight could produce five times more energy per solar panel without batteries, but Feldman cautions that distributed clustering may be a “last 10%” problem that consumes 80% of the development time.

  • Cerebras’s silicon bet is that beating NVIDIA materially requires abandoning GPU-like architecture, because the odds of building a better GPU are “approximately zero.” Its dinner-plate-sized chip places fast memory beside compute to attack AI’s data-movement bottleneck; Feldman says OpenAI workloads run 15-18 times faster than on a GPU.

  • The liquidity debate does not end at an IPO because, as Feldman put it, “more money’s made after IPO than before.” Most early Planet investors retained shares through its public-market re-rating, while Cerebras investors—including Altimeter—were still under lockup and had adopted a six-month “dribble lockup” tied to performance hurdles.

  • Gerstner challenged the idea that Anthropic, OpenAI, and SpaceX’s enormous private valuations are the new normal. Chamath contrasted SpaceX’s prospective scale with historical tech companies that went public at a few billion rather than a few trillion, saying an equivalent post-IPO liftoff would require “quadrillion valuations.” The alternative is an earlier return to public ownership, where “iron sharpens iron” and more investors participate in the upside.

  • 🔗 Original source & video: The IPO Comeback: Why Tech Giants Are Finally Going Public | All-In Liquidity IPO Panel

Listen to full conversation →


Cerebras CEO on the Future of Data Centres, Token Costs & Memory | Should US Companies Sell to China

  • 🗓️ Date2026-05-26 | 🎙️ Show:20VC

AI demand is running ahead of infrastructure, with Cerebras carrying a $25B backlog after models became genuinely useful somewhere in 2025. Structural HBM shortages could persist for years as new fabs require $40B and five years, while Cerebras’s SRAM architecture avoids HBM and CoWoS constraints. Cerebras reports Kimi K2.6 running 6.7x faster than the next-fastest GPU cloud, while its $20B+ OpenAI deal keeps customer concentration in focus.

View Dialogue Notes & Key Takeaways
  • Feldman’s core anti-bubble argument inverts the historical analogy: rail in the 1880s and fiber in the late ’90s were “if we built it, they would come” — infrastructure ahead of demand — while AI is the exact opposite. “We can’t build data centers fast enough to keep up with demand”: Cerebras carries a $25B backlog, and Nvidia, AMD and others have their own. “We’re not building ahead of demand. We’re building behind demand.”

  • The demand inflection has a date: “somewhere in 2025 the models got smart enough to be really useful” — before that, AI “was like cool and then nobody used it.” Now it’s sweeping demographics from his 85-year-old father to his 11-year-old niece, and demand does not peak if AI continues to improve in usefulness — the one bear case he concedes against the “electricity into intelligence” thesis.

  • On the compute deal involving Elon: “They bought down rev gear” — H100s, not B200s, “a generation and a half, maybe two generations behind.” “This was not a great deal. It was a good deal for Elon” — forced action in an exponential, versus Sam Altman’s superpower of believing exponential demand data years out and contracting for power, data centers, and hardware ahead of everyone.

  • Memory is the No. 2 shortage after fab space, and it’s structural: HBM comes only from Samsung, Micron and Hynix, a new fab costs $40B and takes 5 years, so “we’re going to continue to see memory shortages for at least the next several years” if demand holds. Micron is printing 80–85% gross margins — “software gross margins on making memory.” Cerebras sidesteps it entirely: SRAM etched by TSMC, no HBM, no CoWoS, 5nm while 3nm is the oversubscribed node.

  • The speed thesis is absolute: “How big is the market for slow search? It’s zero”— you wouldn’t take $1,000/month to keep slow internet, so there will be zero market for slow inference. Cerebras posted Kimi K2.6 running 6.7x faster than the next-fastest GPU cloud “while one bozo at an analyst firm was on TV saying we couldn’t do it,” and is now digesting a $20B+ OpenAI deal — same concentration critique investors gave him at $1B with G42, one customer-size rung earlier.

  • Nvidia has “funded and backstopped and overallocated to the neo clouds” to create hyperscaler competitors — “a dependence, which is probably not healthy.” Neo clouds buy hardware carrying Nvidia’s 70–80% gross margin before adding their own; Google and Cerebras putting their own silicon in their own data centers don’t, though Google’s full-stack edge is capped by having only one TPU customer: itself.

  • On jobs: “most of the layoffs were AI-washed” — 90–95% of terminations are COVID over-hiring and old productivity gains finally harvested, “none of this is AI yet.” Meanwhile the real enterprise adoption blocker isn’t data cleanliness but lawyers and security (“no credit, no credit, failure, blame”), and Feldman sees no problem with software engineers using $50–100K/year in tokens — hardware engineers’ EDA tools are already closer to 15–20% of salary. At 47M software engineers, that is “$5 trillion just in software engineering token use.”

  • On China, arguing against his own book: leading-edge chips sold to China will be used by its military and its state-backed industry — “there is no debate on that point” — so don’t sell; “I’d like to keep my industrial adversaries more than down rev.” His one free policy wish: give TSMC and Samsung 20 years free of all local ordinances to build US fabs — “fabs are modern pyramids.” The IPO itself was unblocked when “we got a new government” and likely-CFIUS concerns over large customers “disappeared” — the largest semiconductor IPO ever, $185 to $311.

  • 🔗 Original source & video: Cerebras CEO on the Future of Data Centres, Token Costs & Memory | Should US Companies Sell to China

Listen to full conversation →


Google I/O 2026, Karpathy Joins Anthropic, and Cerebras’ $95B IPO | EP #256

  • 🗓️ Date2026-05-21 | 🎙️ Show:Moonshots

Google’s AI comeback is a distribution-and-infrastructure thesis: Gemini reached 900 million monthly users, AI Mode exceeded 1 billion, and projected CapEx is $180–190 billion as Spark and Search embed agents in defaults and commerce. Cerebras’ 68% IPO surge to a stated $95 billion market cap validates wafer-scale inference, but the more-than-$20 billion OpenAI deal and decades-long domestic-fab constraints leave demand and supply to monitor.

View Dialogue Notes & Key Takeaways
  • Google’s comeback is a distribution-and-infrastructure thesis, not a single-model victory. Sundar Pichai said monthly processing rose from 9.7 trillion tokens two years ago to 480 trillion last year and 3.2 quadrillion today, while Gemini reached 900 million monthly users and AI Mode passed 1 billion. CapEx climbed from $31 billion in 2022 to a projected $180–190 billion this year, yet the stock rose—the counterfactual Dave Blundin said nobody would have believed five years ago—and Google is now “disrupting the disruptors.”

  • Google’s technical portfolio is bifurcating between a distinctive multimodal bet and throughput-led models the panel did not mistake for intelligence leadership. Gemini Omni can generate and conversationally edit video from text, photos, video, and audio; Alexander Wissner-Gross called multimodality Google DeepMind’s increasingly “idiosyncratic bet” as US rivals focus elsewhere. Gemini 3.5 Flash, meanwhile, claims four-times-faster output than other frontier models, but Wissner-Gross rated its raw capability “solidly mid” and saw the release as throughput- and tool-use-maxing ahead of Gemini 3.5 Pro.

  • Antigravity 2.0 and Gemini Spark may be fast-following products, but Google’s defaults make them strategically dangerous. Antigravity adopts the agent-first interface already associated with Cursor, while Spark runs continuously in a Google Cloud VM and acts across Gmail, Sheets, and other services. Wissner-Gross’s distribution argument was blunt: Google can be “as good one day later,” integrate with Search, Chrome, and Android, and make billions of users’ first persistent agent a Google agent.

  • Search is becoming a persistent intent-and-commerce layer capable of steering both questions and transactions. Google’s agents can continuously scan for apartments or price changes, while AI suggestions can reshape the query itself; Blundin flagged the “astounding” monetization power for a business protecting roughly $200 billion of 90%-margin revenue. Universal Cart then connects Search, Gemini, YouTube, and Gmail to deal monitoring and checkout—the path Salim Ismail summarized as “intent to agent to transaction,” with the deeper disruption being “getting rid of shopping.”

  • Synthetic abundance makes verification, rather than creation, the scarce asset. SynthID has watermarked more than 100 billion images and videos plus 60,000 years of audio, and OpenAI, Kakao, and ElevenLabs are joining NVIDIA in adopting it. The panel framed this as industry self-regulation moving faster than government and as a transition “from the information age to the verification age,” where provenance becomes infrastructure.

  • Google is extending AI simultaneously into ambient interfaces, scientific research, and rapid entrepreneurship. Its first audio-only intelligent glasses arrive in the fall, trading a visual display for battery life and private spoken assistance while intensifying concerns about recording, backlash, and “the loss of presence in life.” Gemini for Science targets papers, code, hypotheses, weather simulation, and drug discovery; the Build with Gemini XPRIZE gives teams 90 days to address a problem affecting at least 100,000 people, though the presentation cited $2 million while Peter Diamandis later described $3 million in prizes plus $1 million for operations.

  • Andrej Karpathy’s move to Anthropic signals that frontier knowledge and compute are concentrating inside a handful of labs. He will use Claude to accelerate Claude’s own pre-training research after warning that, outside a frontier lab, one’s “judgment fundamentally will start to drift.” The panel read his choice of Anthropic—and his willingness to re-enter a large institution after OpenAI, Tesla, and Eureka Labs—as evidence that independent researchers risk missing the formative inner loop.

  • Cerebras’ 68% IPO surge to a stated $95 billion market cap validated a wafer-scale inference bet that spent years ahead of demand. Andrew Feldman said Cerebras built a dinner-plate chip 58 times larger than any predecessor, filled it with SRAM, and now runs inference 15–20 times faster than GPUs; system sales progressed from 12 in generation one to 300–350 in generation two and “many, many thousands” in generation three. The company subsequently signed a more-than-$20 billion multiyear OpenAI deal, but Feldman warned that domestic fabs remain a decades-long constraint and put Elon Musk’s Terafab ambition in the 15–20-year category.

  • 🔗 Original source & video: Google I/O 2026, Karpathy Joins Anthropic, and Cerebras’ $95B IPO | EP #256

Listen to full conversation →


The Story Behind Cerebras’ $63 Billion IPO with Founder and CEO Andrew Feldman

  • 🗓️ Date2026-05-21 | 🎙️ Show:No Priors

Cerebras says useful AI in 2025 made inference latency a daily constraint, with systems running 15, 18, 20x faster than GPUs. G42’s $1 billion order enabled cluster testing, while an OpenAI agreement Feldman says exceeds $20 billion and an AWS deployment raise the stakes for a 10x manufacturing increase this year. Delivery and governance now matter as much as contrarian architecture.

View Dialogue Notes & Key Takeaways
  • Cerebras’s commercial inflection arrived when models became useful enough in 2025 for inference speed to become a daily-work constraint. Feldman claims its AI computers run inference “15, 18, 20x faster than GPUs” across model sizes and origins. His categorical demand thesis: “How big is the market for slow inference? It’s zero.”

  • That performance rests on a long-running wager that radical gains would require an architecture unlike the GPU, not a minor modification. Cerebras built a 46,000-square-millimeter wafer-scale chip—“the size of a dinner plate”—and spent about $8 million monthly during a 2017-19 stretch when it could not make the design work. It finally yielded the chip in summer 2019.

  • G42’s $1 billion order was the bridge from niche supercomputing customers to hyperscale readiness. It let Cerebras transform its supply chain, deploy and battle-test large clusters, and do training and inference at a scale no internal QA lab could reproduce. Feldman contrasts roughly a dozen first-generation sales and 300 second-generation sales with “tens of thousands” expected for the third.

  • The immediate public-market thesis is delivery against an OpenAI deal Feldman says is north of $20 billion. OpenAI signed after trials showed Cerebras substantially outperforming alternatives; AWS subsequently agreed to deploy its systems in AWS data centers. Cerebras is trying to increase manufacturing 10x this year.

  • The IPO was framed less as an exit than as slightly cheaper capital, audited legitimacy and “corporate adulthood.” The hosts introduced Cerebras at roughly a $60-$63 billion market capitalization with 800-850 employees. Feldman’s differentiation claim is that Cerebras would be the first and only, for a period, AI pure play with 100% of revenue tied to this market—“no gaming, there’s no graphics, there’s no PC.”

  • Cerebras’s own coding adoption shows both AI’s operating leverage and its uneven distribution. Over the past eight months, per-engineer token spending went from below $1,000 to roughly $25,000-$30,000; a small cohort governs eight or 10 agents 24/7 and has moved from 10x to 100x productivity. Feldman includes himself among the rest still “limping along.”

  • Feldman’s larger bet is that low latency will create AI-native businesses, not merely make today’s software faster. His analogy is Netflix: faster internet did not incrementally improve DVD delivery; it helped turn Netflix into a movie studio. Likewise, once companies fundamentally reorganize work around AI, new business models and fundamental jumps in productivity should emerge.

  • 🔗 Original source & video: The Story Behind Cerebras’ $63 Billion IPO with Founder and CEO Andrew Feldman

Listen to full conversation →


Coinbase CEO’s Top 3 Crypto Trends for 2026 + More from Davos!

  • 🗓️ Date2026-01-23 | 🎙️ Show:All-In

Coinbase says U.S. crypto has shifted from regulatory defense to institutional deployment, with five of the top 20 global banks using its infrastructure. Under the GENIUS Act, regulated stablecoins hold 100% reserves in short-term Treasuries, while Coinbase can pass customers “about 100% of the economics” through rewards. B2B cross-border settlement is already driving strong Coinbase Business demand, while trade groups Armstrong believes are trying to revisit the law remain a risk to monitor.

View Dialogue Notes & Key Takeaways
  • Coinbase says the U.S. crypto regime has flipped from attempted extinction to institutional deployment. Brian Armstrong argues the Biden administration tried to “unlawfully kill this industry,” while Trump kept his promise to pursue a U.S. “crypto capital of the world.” Five of the top 20 global banks now use Coinbase infrastructure—including disclosed integrations with JPMorgan and PNC—while BlackRock wants to tokenize every fund.

  • The stablecoin contest is now a fight over Treasury economics, deposits and whether banks can reopen legislation passed four months earlier. Under the GENIUS Act, regulated stablecoins must hold 100% reserves in short-term Treasuries—roughly a 30-day maximum maturity, Armstrong believed—while Coinbase can distribute rewards when customers also trade, make payments or subscribe to Coinbase One. Armstrong says Coinbase passes customers “about 100% of the economics” and that bank trade groups he believes are trying to undo the law represent a “red line.”

  • Armstrong’s three leading crypto trends are the “everything exchange,” prediction markets and stablecoin payments. Equities and other assets are moving on-chain; Coinbase currently works with Kalshi, is talking with Polymarket and could operate its own prediction markets. The clearest stablecoin product-market fit is already B2B cross-border settlement, replacing seven-day transfers and high FX fees; demand for Coinbase Business is strong enough to create an onboarding backlog.

  • Tokenization’s biggest payoff may be cheaper private-market formation and liquidity, provided issuers retain control. Armstrong says private-company tokenization should require the company’s permission, because vesting and illiquidity can retain employees. He expects both fundraising and eventual public listings to move fully on-chain. Coinbase Tokenize targets funds and real estate, while Armstrong frames four billion “unbrokered” adults as the latent market for $100 or $1,000 allocations currently denied access to high-quality assets.

  • Crypto and AI converge when autonomous agents need native wallets and programmable money. Armstrong expects agents to use stablecoins because traditional finance assumes a human behind each product; inside Coinbase, an AI connected to Slack, Google Docs, Salesforce and other systems already surfaces hidden disagreements and audits his time allocation. His preferred mode is “reverse prompting”: asking the system what he should notice or how he could become a better CEO.

  • Cerebras is betting that inference latency, not merely model quality, will determine AI usage and market share. The displayed wafer-scale engine was described on-air as containing 4 trillion transistors and being 56 times larger than a B200, and Feldman says Cerebras aims to collapse multi-stage deep research from minutes to seconds—a “fundamental change in kind,” like broadband turning Netflix from a DVD service into a studio. OpenAI’s announced 750-megawatt Cerebras cloud order makes power delivery, rather than chip count or floor space, the operative unit of capacity.

  • Feldman sees no near-term AI-compute glut, but he does see an 18-month memory digestion and an unresolved geopolitical race. Consumer usage could rise from six or eight queries daily to 100, while every request consumes more inference; simultaneously, inflated 18-month orders have scrambled memory-demand signals and kept memory prices high, with HBM demand adding pressure. China remains behind in high-speed chips but ahead in open models and grid buildout, creating a recursive race where “by getting ahead, you get further ahead.”

  • Gecko Robotics argues that the highest-return AI opportunity is the physical economy, where usable training data barely exists. Jake Loosararian says defense is roughly 30% of Gecko’s business, with Admiral Houston cited as reporting manufacturing-speed improvements as high as 90%, while energy customers use robots to extend asset life and increase output. The roadmap runs from inspection to automated repair and welding, with skilled humans supervising fleets; generic bricklaying may arrive in roughly three years, but industrial autonomy requires proprietary data gathered inside refineries, shipyards and power plants.

  • 🔗 Original source & video: Coinbase CEO’s Top 3 Crypto Trends for 2026 + More from Davos!

Listen to full conversation →


Cerebras CEO, Andrew Feldman on Why Raise $1BN and Delay the IPO & Why NVIDIA’s Worried About Growth

  • 🗓️ Date2025-10-06 | 🎙️ Show:20VC

Cerebras raised $1bn at a record valuation because it had margins while rivals had negative ones, despite still intending to go public; UAE orders may have represented 75-80% of revenue and consumed manufacturing capacity. Feldman argues two-year chip depreciation is empirically wrong and Nvidia’s pre-announcements and OpenAI investment signal defensive demand capture, while power sits in the wrong places; investors should watch utilization, customer concentration, permitting, and whether AI adoption delivers productivity.

View Dialogue Notes & Key Takeaways
  • Cerebras raised $1bn — the largest round ever in its category, at the highest valuation — led by Fidelity (“the Oxford or Cambridge of investing”) with Tiger Global, Valor and 1789, and Feldman says Cerebras still intends to go public: “We still have every intention of going public.” Feldman says the round was possible because Cerebras had margins while rivals shopping for money had negative ones — and he thinks the S-1 said the UAE accounted for about 75-80% of revenue, perhaps in H1 2024, with orders so large they “consumed all our manufacturing capacity.”

  • Nvidia is showing a worried giant’s tells: “use your balance sheet more and your technology less” — buying business rather than winning it (as Cisco did from 1992-2001), “predatory pre-announcing” B300s before anyone can get B200s, and silence on field failure rates “which are massive.” The $100bn OpenAI investment “was designed for nobody to understand it… it’s just not an analyzable thing” beyond Nvidia locking up a slice of OpenAI’s demand.

  • Two-year chip depreciation is “empirically wrong”: H100s are still earning past two years, A100s are at three-to-four and “could be as long as five or six.” The real depreciation variable is generational improvement — and apples-to-apples (8-bit to 8-bit) Feldman estimates 2-2.5x per meaningful generation, because memory bandwidth, not flops, is the binding constraint for inference on GPU architectures.

  • Demand is unknowable even to buyers — customers ask Cerebras for “between five and 40 million queries per second,” with Harry highlighting the uncertainty as an order of magnitude — so read megadeals as options on the future, and read “up to $100 billion over five years” as “the great CYA word in marketing history.” Chance he’s still underestimating demand: “100%. I’ve been wrong.”

  • On Mag7 concentration, the risk isn’t valuation (Nvidia at $4.5T: “maybe it’s too low”) — it’s mispriced diversification: the S&P “is not an index of the global economy — it’s 30% or 50% seven companies,” so holders carry sector risk they never signed up for. “Risk comes in financial markets where people fundamentally underestimate risk.”

  • On energy, the scarcity narrative is “strictly wrong”: “We have plenty of power. It’s in the wrong places” — West Texas gas and upstate New York hydro sit far from people and fiber, nuclear is reasonable but not unavoidable, and America’s patchwork permitting (a local fire ordinance set Samsung’s Texas fab back 8-10 months) is the real handicap versus China’s central planning.

  • Contrarian diffusion calls: Groq’s Jonathan Ross’s five-year AI labor-shortage prediction is “absolutely wrong — that might be true in 15 years”; AlphaFold won Nobels but “name a drug that’s resulting from it. Not one.” Productivity jumps only when society reorganizes around AI — the electricity-and-the-dynamo lesson — not when AI just replaces Google.

  • 🔗 Original source & video: Cerebras CEO, Andrew Feldman on Why Raise $1BN and Delay the IPO & Why NVIDIA’s Worried About Growth

Listen to full conversation →


20VC: The Daily Deal: Coreweave IPO | Scale Hits $25BN on $2BN EOY Revenue | Sequoia’s 25x Return on Wiz | Tech Stocks Tank with Tariffs | Cursor: Defensible or Dangerous Example of Lost Moats in Tech

  • 🗓️ Date2025-04-03 | 🎙️ Show:20VC

AI is creating a golden age of software and extraordinary revenue growth, but thin wrappers, dilution, and uncertain recurring revenue are exposing the limits of the old scale-equals-durability model. Moveworks’ $2.85B sale shows the value of enterprise distribution and workflow integration, while CoreWeave’s IPO, a possible $2B put obligation, and constrained M&A make execution and liquidity the next catalysts.

View Dialogue Notes & Key Takeaways
  • Jason Lemkin’s central call is that AI is opening a “golden age of software” while destroying the old assumption that scaled revenue is durable. If software expands from 2% to 4% of GDP, he sees 100 or more decacorns. Harry’s framing was that every VC has five or ten unicorns that are no longer unicorns, now growing in the single digits or teens; Lemkin expects late-stage slowdowns to become commonplace by 2026 and those stranded assets to become less attractive to acquire.

  • Hot AI rounds can look cheap at roughly 20x forward ARR only because investors extrapolate extraordinary slopes while remaining uncertain whether the “R” is recurring and durable. Lemkin says markups still grade VCs and some LPs, while Harry notes that entering at $200M ARR can shorten duration; the counterweights are dilution that Lemkin estimates might approach 10% annually and a venture cash cycle stretching toward 20 years.

  • “Triple-triple, double-double” remains elite operating performance but is currently “silver in a gold rush” to perhaps 80% of B2B investors. Lemkin’s prescription is a year-long, relationship-led process with roughly 50 investors and honest monthly updates: “It only takes one.” Founders without Cursor-like heat should avoid artificial Friday deadlines and may be better off accepting a valuation 30% lower if that gets the round closed.

  • AI demand is producing real revenue faster than technical substance, while investors remain “addicted to top-line growth.” Lemkin has seen companies reach $2M in 60 days or eight figures inside a year with humans running prompts behind thin wrappers; yet AI efficiency does not necessarily reduce capital needs, because RevenueCat took a 2x productivity gain and “plowed it all into new hiring.”

  • Moveworks accepted ServiceNow’s $2.85B offer after years of deep integration and accelerating demand made ServiceNow’s distribution especially compelling. Bhavin Shah said 250 of 350 customers already used ServiceNow, while Moveworks had five million users against ServiceNow’s 150M-plus; he concluded that Moveworks could not match the market’s scale and speed independently. The strongest ROI was not saving an employee three hours, but automating core workflows across systems such as SAP, Workday, Salesforce, Concur and Jira.

  • Lemkin expects an IPO and M&A “gold rush” within roughly 18 months, but sees little evidence that PE will rescue slow-growth unicorns. He expects large private companies such as Stripe, Figma, Chime and Canva may list, yet says cash-flow-positive assets at $20M, $50M or even $100M are not attracting the “tire kicking” he once saw; PE instead appears to be combining holdings into “Frankensteins” such as SalesLoft with Drift and Gainsight with Skilljar.

  • CoreWeave getting public was a major entrepreneurial achievement, but its real scorecard begins on days 180, 365 and 450. Andrew Feldman called an IPO “the beginning of adulthood”; Lemkin nevertheless worries that, as he understands the structure, last-round investors can put back almost $2B of stock if shares fail to trade 70% above the IPO price within two years. That deadline could invite shorts and sacrifice long-term decision-making to a fixed date.

  • Feldman’s warning on fashionable hardware and defense investing is that accumulated experience still matters. A chip tape-out can consume $20M-$30M in non-recurring engineering costs—and a bug means paying again—while defense requires trusted relationships, cleared personnel, specialized facilities and patience with “Bible-sized contracts.” Commercial technology can still reshape warfare, but procurement and cost-plus incentives remain the constraint.

  • 🔗 Original source & video: 20VC: The Daily Deal: Coreweave IPO | Scale Hits $25BN on $2BN EOY Revenue | Sequoia’s 25x Return on Wiz | Tech Stocks Tank with Tariffs | Cursor: Defensible or Dangerous Example of Lost Moats in Tech

Listen to full conversation →


Andrew Feldman, Cerebras Co-Founder and CEO: The AI Chip Wars & The Plan to Break Nvidia’s Dominance

  • 🗓️ Date2025-03-24 | 🎙️ Show:20VC

Cerebras argues that GPU off-chip HBM is a fundamental inference bottleneck, while wafer-scale SRAM could challenge NVIDIA’s near-total market share. Feldman sees inference demand exceeding 100 times today’s market within five years, with chip providers potentially worth more than model companies, though G42’s 87% revenue concentration remains a key risk.

View Dialogue Notes & Key Takeaways
  • Feldman’s core attack on NVIDIA: the GPU’s off-chip HBM memory is a fundamental architectural limitation for generative inference — serving one word from a 70B-parameter model means moving ~140GB of weights from memory to compute, again for every next word. Wafer scale let Cerebras use fast SRAM at capacity, and “what used to be their advantage is now weakness… it can be beaten and I think they know it.” His market-structure call: NVIDIA goes from “approximately all” of the market today to 50-60% share in five years — between Uber’s 90/5 and cloud’s even split.

  • The CUDA moat is “not real at all” in inference — “you can move from OpenAI on an Nvidia GPU to Cerebras… with 10 keystrokes.” The real, rarely discussed moat is market-share leadership itself: Intel made “nearly a decade of catastrophic decisions” and still holds ~75-80% of x86. “I can make a bunch of bad decisions for a decade and only lose 20% share — the moat was just unbelievable.”

  • Scaling-law gains continue — Feldman flatly rejects the “far along on compute, algorithms and data” consensus: “I think they’re wrong. I think we are early in all of them.” A GPU doing inference runs at 5-7% utilization — “95 or 93% wasted” — OpenAI’s o1 shows inference scaling laws “fully functional,” and we won’t be as dependent on Transformers in 3-5 years (“100%”). In five years training data is “almost all synthetic.”

  • The inference market equation: users × frequency × compute-per-use — all three growing simultaneously, a rare condition. AI flipped from “novelty” to useful in Q4 2024, the market in five years is “way over 100 times bigger,” and faster/cheaper only expands it: “no examples in compute in 50 years in which by making things cheaper, faster, the market got smaller.”

  • Chip providers will be worth more than model providers on a 5-year view. Today’s model valuations are option pricing — “uncertainty is a friend of the value of the option” — but Buffett’s weighing machine eventually kicks in. On models generally: competing on release cadence four months ahead of rivals carries “not a lot of value”; staying top-decile for years does.

  • Cerebras is cash-flow positive where peers hemorrhage cash — “gross margins were a measure of your technical differentiation… if you’re running a negative gross margin business, you’re selling commodity.” The flip side: G42 is 87% of revenue (a deal estimated north of $1bn), which Feldman defends as a learned muscle: “the way you catch three large customers is to catch one first,” with several relationships targeted in 24 months.

  • On China: Cerebras refused to sell — the deal “wouldn’t be used for good” (facial recognition of minorities, military) and failed his mother test — yet he thinks the US underestimates China “100%,” export controls may not be “a tractable problem,” and the attempt to cut off EDA tools just spawned US-VC-backed EDA startups in Shenzhen.

  • 🔗 Original source & video: Andrew Feldman, Cerebras Co-Founder and CEO: The AI Chip Wars & The Plan to Break Nvidia’s Dominance

Listen to full conversation →