Dylan Patel — The single biggest bottleneck to scaling AI compute
Dylan Patel — The single biggest bottleneck to scaling AI compute
Summary
- Anthropic’s compute conservatism is now a tax. Dylan’s math: adding ~$6B revenue/month implies ~$40B of inference compute at current gross margins — four incremental gigawatts just to grow — so Anthropic must scrape to 5-6GW by year-end via Bedrock/Vertex/Foundry revenue shares and last-minute capacity, while Dario “screwed the pooch” versus OpenAI’s “let’s just sign these crazy fucking deals.” H100 deals are printing at $2.40/hr for 2-3 years against a $1.40 all-in build cost.
- The GPU depreciation bears have it backwards: because chips are supply-constrained, an H100 is priced off the value GPT-5.4 can extract (TAM “north of a hundred billion”), not off Blackwell’s spec sheet — “an H100 is worth more today than it was three years ago.” Michael Burry’s sub-3-year depreciation lens only works “if you could build infinite Rubin,” which you can’t.
- The 2030 binding constraint is ASML, and the ceiling is roughly ~200GW/year under full AI allocation. One gigawatt of Rubin needs ~2M EUV passes = 3.5 EUV tools; the installed base plus 70→80→~100 tools/year gets to ~700 tools by decade-end, i.e. ~200GW of AI chips annually. Sam Altman’s gigawatt-a-week is “completely compatible” — it’s only ~25% share — but everyone (Sam + Elon’s 100GW/yr in space + Anthropic + Google) cannot simultaneously get what they want.
- The supply chain is structurally under-building because nobody below the labs is AGI-pilled: labs need X, Nvidia builds X−1, downstream builds “X ÷ 2.” The tradeable analog to the 2023-25 turbine-deposit trade: pay ASML for forward EUV options and resell “once other people realize everything is fucked” — except ASML, “one of the most generous companies in the world,” probably won’t agree.
- Memory is the nearest crunch: ~30% of Big Tech’s 2026 CapEx goes to memory, prices double again, and smartphone volumes fall from 1.1B toward 500-600M as Xiaomi/Oppo halve low-end lines. There’s no DDR escape hatch — HBM4 delivers ~2.5TB/s per stack versus ~64-128GB/s of DDR on the same chip shoreline — and new fabs don’t land until late 2027/2028.
- Google fumbled the TPU dislocation: Anthropic’s ex-Google compute team locked up a million Ironwood v7s before Gemini’s ARR hit $5B in Q4, DeepMind people said “this is insane, why did we do this,” and by the time leadership woke up TSMC replied “sorry guys, we’re sold out.” Google has since gone “absurdly AGI-pilled” — buying an energy company, turbine deposits, powered land.
- “Fast timelines, the US wins; long timelines, China wins.” China gets fully indigenized DUV by 2030 “for sure” and working-but-not-volume EUV; meanwhile US labs hit ~10GW each by end of next year and distillation gets harder as the product shifts from tokens to opaque automated work. If returns are middling and AI takes longer to reach certain capability levels, China’s vertical supply chain scales past the West — and Huawei, unbanned, “would have already eclipsed Apple as the biggest TSMC customer.”
- Power and space are not the story. Sixteen-plus gas-adjacent vendors, ship engines, Bloom fuel cells, and batteries or peaker plants unlocking ~20% of the terawatt US grid make power solvable at 2x cost (“the Hopper that was $1.40 is now $1.50 in cost. I don’t care”); space data centers lose 10% of useful GPU life to deployment lag and are a next-decade 10X, not this one. The real tail risk stays Taiwan: lose it and incremental capacity collapses to ~10-20GW/year.
Deep dive
1. $600B of CapEx buys ‘27-‘29 — and Anthropic is buying compute in a pinch
- Dylan’s decomposition of the big four’s $600B CapEx (~$1T across the supply chain): only a portion pays for the ~20GW of US capacity landing this year. Google’s $180B includes turbine deposits for ‘28 and ‘29, data center construction for ‘27, and power-purchase down payments — setup spend “so they can set up this super fast scaling.”
- The Anthropic bind, run forward: ~$6B of revenue added per month implies ~$60B over ten months, which at last-reported gross margins means ~$40B of inference compute — four gigawatts at ~$10B/GW rental — just to grow revenue, training fleet flat. Dylan thinks they scrape to “five or six gigawatts” by year-end via Bedrock/Vertex/Foundry revenue shares, with OpenAI “a little bit higher.”
- The contrast Dylan draws: Dario was deliberately conservative — “he’s screwed the pooch compared to OpenAI, whose approach was, ‘Let’s just sign these crazy fucking deals’” — while OpenAI’s aggressive approach extended to random counterparties like SoftBank Energy, “which has never built a data center in its life.” The Anthropic meme, offered with a grin: “they have commitment issues and are sort of polyamorous.”
- The four-month whiplash worth remembering: everyone refused OpenAI deals because “you don’t have the money,” Oracle and CoreWeave tanked, credit markets panicked — then the raise landed and “now everyone’s saying, ‘OpenAI, we believed you the whole time.’” Anthropic, the responsible one, is the party now paying spot markups.
2. The H100 is worth more today than three years ago
- Dylan’s two lenses on depreciation. The mechanical TCO lens: an H100 costs $1.40/hr across five years; $2/hr contracts are ~35% gross margin; if you could build infinite Rubin, Hopper falls to $1/hr in ‘26 and $0.70 in ‘27 as each generation triples performance for 1.5-2x price. That’s the Burry case — and it’s the wrong lens.
- The right lens: chips are priced off the value derivable today. GPT-5.4 is sparser, cheaper to serve, and an H100 pushes more tokens of a better model than it ever did of GPT-4 — GPT-4’s token TAM was maybe tens of billions, GPT-5.4’s “probably north of a hundred billion.” Result: “an H100 is worth more today than it was three years ago. That’s crazy.” Signed at $2.40 two years into its life, Hopper’s margins have expanded, not decayed.
- The AGI extension: an H100’s ~1e15 flops is in brain range, and “if we had actual humans on a server, the value of an H100 is such that it can repay itself in the course of a couple of months.” (Dwarkesh’s petabyte-brain objection drew the episode’s best heckle: “Name a petabyte of ones and zeros, bro. Name me a string.”)
- Dwarkesh’s Alchian-Allen framing — worth keeping: a fixed increase in compute cost compresses the price ratio between frontier and mid models, pushing everyone toward the best model. Dylan buys it: “we just see all the volumes are on the best models today, all the revenue is on the best models today.”
3. Margin flows upstream to whoever committed early — and labs must destroy demand
- In a compute-limited world, early five-year committers “locked in a humongous margin advantage” — but the flex is limited: CoreWeave’s average term is 3+ years on 98%+ of compute, so incremental capacity is where the repricing happens. And incremental capacity is enormous: Meta alone is adding this year as much as its entire 2022 fleet for WhatsApp, Instagram, Facebook and AI combined; OpenAI went 600MW → 2GW → 6+ → 12 next year.
- Upstream holds the cards: Nvidia has $90B of long-term contracts and is negotiating three-year memory deals; memory vendors “are going to double or triple price again”; TSMC isn’t raising, ASML isn’t raising — “until TSMC or ASML break out and say, ‘No, we’re going to charge a lot more.’”
- Model-vendor margins go up this year for an unhappy reason: “because they’re so capacity constrained, they have to destroy demand. There’s no way Anthropic can continue at the current pace without destroying demand” — visible already in Claude’s reliability.
4. Google sold a million TPUs to Anthropic before waking up
- Dwarkesh cites Dylan’s numbers putting Nvidia at ~70%+ of N3 by ‘27: it signed non-cancelable, non-returnable orders way earlier than Google or Amazon, checked PCB and memory capacity across the chain, and was “way more AGI-pilled than Google was in Q3 of last year.” TSMC’s own bias cuts the same way — it prefers allocating to Graviton CPUs over Trainium because CPU demand looks like “stable, long-term growth,” and Apple is sliding toward being “just like any old customer” as A16’s first customer is AI, not Apple.
- The Ironwood story as Dylan reconstructs it from wafer data: Anthropic’s compute team — both ex-Google — “saw this dislocation” and negotiated before Google realized; over six weeks in early Q3, TPU capacity requests jumped multiple times, so suddenly Google had to explain the increase to TSMC. DeepMind people: “This is insane. Why did we do this?”
- Then Nano Banana and Gemini 3 hit, Gemini reached $5B ARR in Q4 from next to nothing in Q1-Q3, and leadership went back for more. TSMC: “Sorry guys, we’re sold out. We can maybe get 5-10% more for 2026.” Dylan’s verdict, hedged as his own narrative from supply-chain data: “it’s pretty clear to me that Google screwed up.”
- Since then Google has “gotten absurdly AGI-pilled” — bought an energy company, put deposits on turbines, is buying “a ridiculous percentage of powered land.” Asked how many gigawatts they’ll have by end of next year: “Buy my data.”
5. The bottleneck migrates to chips — and the EUV math caps AI at ~200GW/year
- The bottleneck sequence — CoWoS, then power, then data centers — were all short-lead-time problems (Amazon builds data centers in eight months; fabs take 2-3 years). The mobile/PC fungibility that fed AI so far is exhausted: Nvidia is now the largest customer at both TSMC and SK Hynix, and “it’s sort of impossible for the sliding of resources… to shift any more.”
- The load-bearing arithmetic: a gigawatt of Rubin needs ~55,000 3nm wafers, ~6,000 5nm, ~170,000 DRAM wafers; at ~20 EUV layers of a 3nm wafer’s ~70 masks, that’s ~2M EUV passes per gigawatt. At 75 wafers/hour and 90% uptime, 3.5 EUV tools satisfy one gigawatt — $1.2B of tools underwriting a $50B data center.
- ASML makes ~70 tools this year, 80 next, “a little bit over 100 by the end of the decade” even under aggressive expansion. Stacked on an installed base of 250-300, that’s ~700 tools by 2030 → ~200GW/year of AI chips if fully allocated to AI. Sam’s gigawatt-a-week (52GW/yr) is “completely compatible… he’s only taking 25% share then” — roughly his Blackwell share already. The problem is everyone wants that: Elon wants 100GW/yr in space, Anthropic and Google similar.
- The margin telescope Dylan flags: Nvidia converts a small fraction of TSMC’s $100B of three-year CapEx into $160B of annualized earnings, and ASML sits below both turning ~$1B of machines into a gigawatt.
6. ASML is the most complicated machine humans make — run by the least greedy company
- Why no YOLO expansion: the supply chain “has lived through the booms and busts” and isn’t AGI-pilled — “no one really sees demand for 200 gigawatts a year.” Dylan’s whip-lag formulation: labs know they need X, “Nvidia is not quite as AGI-pilled. They’re building X − 1. You go down the supply chain, everyone’s doing X − 1. In some cases, they’re doing X ÷ 2.” ASML thinks “AI means we have to go from 60 to 100.”
- The artisanal reality: Cymer’s source hits tin droplets three times with a laser to coax out 13.5nm light; Zeiss makes 18 mirror-lenses of perfectly layered molybdenum/ruthenium; the reticle stage moves at nine Gs opposite the wafer stage while holding ~3nm layer-to-layer overlay; the tool is assembled in Eindhoven, disassembled onto multiple planes, and rebuilt on site over months. Supply chain: 10,000+ suppliers, none trainable “in the snap of a finger.”
- ASML “is maybe one of the most generous companies in the world” — a monopoly linchpin that “has never raised the price more than they’ve increased the capability of the tool.” Leopold’s view, per Dylan: “Let’s have the price go up. Because they can.”
- The trade Dylan sketches, echoing the 2023-25 turbine-deposit playbook: “I’ll pay you a billion dollars. You give me the right to purchase ten EUV tools two years from now,” then sell the option “once other people realize everything is fucked and we don’t have enough capacity.” The catch, twice-stated: “will ASML even agree to this? I don’t think so.”
7. You can’t just fall back to 7nm — the ladder of interconnect says no
- Dwarkesh’s fallback: if EUV binds, re-ramp 7nm like China. Dylan’s rebuttal starts with numerics (A100 was designed for FP16/BF16, Rubin for FP4/FP6 — the FLOPS comparison isn’t fair) but the real gate is the communication ladder: on-chip data moves at tens-to-hundreds of TB/s, within-rack ~1TB/s, across racks ~100GB/s — and frontier models run on hundreds of chips (DeepSeek serves production on 160 GPUs).
- The killer datapoint: at 100 tokens/second inference on DeepSeek and Kimi K2.5, Hopper vs Blackwell is “on the order of 20x — not 2x or 3x like the FLOPS performance difference indicates, even though those are on the same process node.”
- Packaging can stretch old nodes — Tesla’s Dojo put 25 chips on a wafer and “is still probably the best chip for running convolutional neural networks,” just shaped wrong for transformers; Huawei’s Ascend 910C/D went from one die to two — but “anything you do on 7nm, you can also probably do on 3nm,” so the leading edge keeps its lead.
- Related system logic Dylan lays out: scale-up domains differ (Nvidia NVL72 all-to-all at TB/s, Google torus pods of 8-9k chips, all three converging on dragonfly), but the deeper reason parameter scaling stalled is RL economics — smaller models do rollouts faster and get deployed into research sooner, and “this compounding effect of being able to do research faster and faster is potentially a faster takeoff.” Google runs the largest production model (Gemini Pro, bigger than GPT-5.4 or Opus) only because its unipolar TPU fleet lets it.
8. Fast timelines, the US wins; long timelines, China wins
- China’s tooling trajectory, as hedged: fully indigenized DUV by 2030 “for sure”; EUV “working tools” but not volume — “there’s having it work, and then there’s production hell,” which took ASML five to seven years. His shot-in-the-dark estimate: China at ~100 DUV tools/year vs ASML’s hundreds.
- Why the US pulls away on fast timelines: distillation from American models gets harder as the product shifts from tokens with visible reasoning chains to automated white-collar work with hidden thinking; OpenAI and Anthropic both hit ~10GW by end of next year while “China is not scaling their AI lab compute nearly as fast.”
- The ROIC case that this is already a takeoff: Anthropic added $4B of revenue in January, $6B in February; its $20B ARR at sub-50% margins implies ~$13-14B of rental compute sitting on ~$50B of someone’s CapEx. “It’s not like we’re talking about a Dyson sphere by X date — the revenue is compounding at such a rate that it does affect economic growth.” The flip side, stated fairly: if Wall Street’s bears are right and returns are middling, China’s fully vertical indigenized chain “is able to scale past us.” Dwarkesh’s pushback: you don’t need to believe in AGI for the US-win timelines.
- The Huawei counterfactual — Dylan’s spiciest call: Huawei is “arguably the only company in the world that has all the legs” (cracked software, networking, AI talent, own fabs, own token end-market), was first to a 7nm AI chip — two months before the TPU, four before A100 — and “if in 2019 Huawei was not banned from using TSMC, Huawei would have already eclipsed Apple as the biggest TSMC customer.”
9. The memory crunch: 30% of CapEx, dying smartphones, and no DDR escape hatch
- Dwarkesh’s escape hatch — ditch HBM for commodity DRAM and ship “Claude Slow” — fails on shoreline physics: an HBM4 stack pushes 2048 bits at 10 giga-transfers = ~2.5TB/s per ~13mm of chip edge, versus ~64-128GB/s for DDR in the same shoreline. The metric is bandwidth per wafer, not bits per wafer — and behaviorally, “no one actually wants to use a slow model”: agentic hours become a day.
- The scale of it: ~30% of Big Tech’s 2026 CapEx goes to memory. Consumer pain is the release valve — an iPhone’s 12GB goes from ~$50 to ~$150 of DRAM, NAND moves too, so “the end consumer is paying $250 more for an iPhone.” Smartphone volumes: 1.4B historically → 1.1B now → maybe 800M this year, 500-600M next, with Xiaomi and Oppo already halving low-end lines. “It’s going to be even worse when memory prices double again.”
- Why supply can’t respond: vendors lost money on memory in 2023 and built no fabs. SemiAnalysis “was banging on the drums” in 2024 — reasoning → long context → KV cache → memory — but pricing took a year to reflect it and fabs take two more, so meaningful new capacity lands late 2027/2028. Interim scramble: Micron buying a lagging-edge Taiwanese fab.
- On Elon’s million-wafer TeraFab: he can probably build the clean room (“all of the air in the fab gets replaced every three seconds”) in a year or two, but his delete-things, it-can-be-dirty mindset is “100% not right” — and process development, the real moat, “has a lot of built-up knowledge.” The 3D-DRAM hope still runs on EUV, but bits-per-pass “would go up drastically” — a genuine late-decade disruptor requiring massive fab retooling.
10. Power is solvable, space is next decade, Taiwan is the real cliff
- Power stopped being the binding constraint because “humans and capitalism are far more effective” than three turbine makers: SemiAnalysis tracks 16+ gas-power vendors with hundreds of gigawatts of orders — aeroderivatives (Boom Supersonic with Crusoe), reciprocating engines, ship engines (Nebius powering a Microsoft data center in New Jersey), Bloom fuel cells, solar-plus-battery. Batteries or peaker plants absorbing the peak could unlock ~20% of the terawatt-scale US grid; ~half of decade-end additions go behind the meter. Cost doubling doesn’t matter: “the Hopper that was $1.40 is now $1.50 in cost. I don’t care.”
- Labor is “a humongous constraint” but a solvable one: Abilene’s 1.2GW site peaked at ~5,000 workers, implying ~400,000 people per 100GW — answered by importing skilled electricians and modularizing in Asian factories: megawatt-class Kyber racks, whole pre-integrated rows, “unfortunately for America.”
- On space data centers, Dylan’s skepticism is operational, not physical: ~15% of deployed Blackwells get RMA’d, and testing, deconstructing, launching, and recommissioning adds ~six months — “10% of your cluster’s useful life” at the moment compute is most valuable — plus unreliable space lasers replacing million-unit-volume transceivers. “Elon doesn’t win by doing 20% gains… Elon wins when he swings for the fences and does 10X gains” — and space becomes that 10X “but that’s not this decade.” He reads Elon’s Samsung-in-Texas robot-chip deal the same way: a hedge because “he thinks Taiwan risk is huge.”
- The Taiwan endgame: airlifting every TSMC process engineer doesn’t save you, because the tools themselves need Taiwanese chips — “a snake eating its own tail meme.” Lose Taiwan and incremental capacity collapses from hundreds of gigawatts a year to “maybe 10 gigawatts across Intel and Samsung, or 20… It’s nothing” — with China suddenly holding the most verticalized chain left standing.