Pioneers Insight Method Research Author
Watts, Wafers, and the Future of AI Infra | Gavin Baker
Back to Episodes

Watts, Wafers, and the Future of AI Infra | Gavin Baker

Summary

  • March/April was a lean-in drawdown, not a take-your-medicine drawdown: while the Nasdaq sold off, Anthropic added $11B of ARR in a single month — equal to the combined businesses Palantir, Snowflake and Databricks took a decade to build. “Nothing like that has ever happened in the history of capitalism,” and tech simultaneously got as cheap versus the market as at any point in the last 10 years. “All you had to do in March was simply observe what was happening to Anthropic.”
  • Anthropic at a rumored $900B on $50B ARR is cheaper than it looks: Baker argues compute constraints and a degraded Claude (Opus generating ~70% fewer tokens per question) mask an “unconstrained run-rate revenue” of $100–200B — so you may be paying ~5x “URR.” They could raise at “at least a 100% premium” but are following Elon’s lesson: “never being greedy on valuation” buys 20–30 years of capital access.
  • The single bubble tell is TSMC’s capacity decisions. If TSMC gave Jensen what he wants, Nvidia could sell “$2 trillion of GPUs in ‘26 or ‘27. Maybe 2½. Maybe 3” — and that’s overbuild territory. History (Carlota Perez, canals, railroads, 2000) says expect a bubble; the offsets are a build-out still funded from operating cash flow and GPUs at 100% utilization vs. 99% dark fiber. Watch for Intel or Samsung breaking discipline: “they’re not going to stay disciplined. They will break.”
  • Orbital compute is “racks in space,” not Pentagon-sized data centers — a Blackwell rack is 3,000 lb and 100kW; Starlink V3 is expected to operate at 20kW, and SpaceX is targeting 100–120kW satellites linked by lasers through vacuum. Watts shortage alleviates ‘27–‘28, then orbital solves it — a direct warning to terrestrial power/cooling suppliers whose capacity ramps land “just as all of the silly skeptics start to understand that orbital compute is very real.”
  • Frontier tokens are capturing the overwhelming majority of model-layer value — Gemini 3.1 Pro went from “mind-blowing” to “intolerable” — and the Pareto frontier flipped: Google dominated 9 months ago; now it’s Anthropic and OpenAI, with Grok 4.3 on the frontier and Gemini clinging on, likely “subsidizing out of pride.” Meanwhile AI’s shift “from all you can eat to pay by the drink” (his telecom-analyst analogy) is probably why OpenAI and Anthropic should exceed well over $200 ARR this year.
  • Prefill/decode disaggregation extends GPU lives to 10–15 years — put Cerebras or Groq LPUs in front of Hoppers and Amperes and run them “until it melts” — which “may single-handedly save private credit” and, by cutting GPU financing from CoreWeave’s low-7s toward 5–6%, mathematically lowers the cost of the whole build-out.
  • Cross-sectional valuations “flat out do not make sense”: semi-cap at 40x next quarter’s annualized earnings vs. DRAM at mid-single-digits (last cycle’s peak gap was ~5 vs. 12); the lowest-quality, highest-cost “sellers of shortage” are mooning on X-driven flows while quality lags. AI intra-correlations blew apart in January — you can no longer hedge memory with semi-cap — and miscategorized names like Astera (a switch company stuck in “copper loser” baskets) are the opportunity. “I just wish there were more AI bears.”
  • The three questions that decide everything: does the frontier-token premium persist, does the bitter lesson survive contact with ASI (“who knows if the bitter lesson holds for 400-IQ models”), and when does continual learning arrive — “if we get that, then we have a really fast takeoff.” Plus the dark coda: rising political violence around AI, and “the machine gun is here… if we do not all become masters of the machine gun, we’re going to get mastered.”

Deep dive

1. March was pent-up alpha — the most extraordinary month in the history of capitalism

  • Baker’s taxonomy: there are drawdowns where “your hypothesis was invalidated and you have to take your medicine,” and drawdowns where you profoundly disagree with the price action and can “build pent-up alpha.” March was the second kind — the Nasdaq sold off into what he calls the most extraordinary moment in the history of American business.
  • The exhibit: Anthropic added $11B of ARR in one month. Palantir, Snowflake and Databricks — arguably the three highest-profile SaaS companies founded in the last 10–12 years — spent a decade and tens of thousands of employees building their businesses. “Anthropic added their combined businesses in 1 month… nothing like that has ever happened in the history of capitalism.” Krishna’s stat on the show — a garbled “500% in DR” statistic (as heard) — compounds to insanity over three years.
  • The DeepSeek rhyme: in ‘25 the paper published seven days before “DeepSeek Monday,” and by the time stocks imploded it was already “super clear this was going to be the most positive thing that had ever happened to compute demand” — DRAM going vertical, Asian AWS availability-zone prices doubling. That trade took some work; this March “all you had to do was simply observe what was happening to Anthropic.”
  • His contrarian macro take: the Strait of Hormuz closing is “actually relatively awesome for America” — natural gas on Bloomberg fell 20% while gas in Asia and Europe doubled or tripled, improving America’s relative manufacturing competitiveness overnight, which is exactly what the administration cares about. Unlike the ’70s, the US is now the largest producer and exporter of oil and gas — which made it easier to stay focused on tech at its cheapest relative valuation in 10 years.

2. Anthropic is ~5x “URR” — and the Elon lesson on never pushing valuation

  • OpenAI and Anthropic are “pretty different animals from a capital efficiency perspective”: Anthropic has burned maybe 80% less to reach roughly similar revenue scale, implying structurally different ROICs, though he credits Sarah Friar as “one of the most exceptional CFOs” and notes OpenAI’s aggression in securing compute “really paid.”
  • His new metric: URR — unconstrained run-rate revenue. Anthropic has “clearly degraded the intelligence of Claude” — analysis shows even Opus generating 70% fewer tokens for the same question, and “token quantity equals quality of answer” — so with all the compute it wanted, Anthropic would be doing $100B, $150B, “maybe 200.” At $900B for $50B of ARR, “you might be buying it at more like five times URR.”
  • Why not raise $100B at $3T? Because Elon’s superpower — being able to raise as much capital as he wants, whenever — was earned by “just never being greedy on valuation. Never pushing valuation. Just that simple.” His friend Antonio’s point: SpaceX compounded low-30s percent a year for a decade. Anthropic could raise at “at least a 100% premium to this rumored latest mark” — of course. This approach can create benefits lasting 20–30 years.

3. Watts: zoning is the new gate, and orbital compute is racks in space

  • Capitalism solves the watts shortage “absent big regulatory or political blowback, which I think is a real possibility.” The head of data-center infra at a big PE firm: “It used to be energy and chips were our biggest gating factors. Now it’s zoning and approval.” Companies may wait until after the midterms to act on possible workforce reductions — “nobody wants to be a piñata.” Turbine constraints are real (two such machines can cast the big blades; the West hasn’t built one in 80 years), but he expects the shortage to begin alleviating ‘27–‘28.
  • Then the reframe he insists on: orbital compute is not a Pentagon-sized building in space. “It’s racks in space.” A Blackwell rack is 3,000 lb, 8ft × 4ft × 3ft, 100kW; the satellite is roughly that, with ~500-ft solar wings in sun-synchronous orbit and a radiator trailing hundreds of feet. Racks link via lasers through vacuum — already on every Starlink — into a virtual data center. In space you optimize for weight, not size; on Earth it’s “copper when you can, optics when you must.”
  • The credibility argument: SpaceX operates 98–99% of all satellites in orbit and the largest data center on Earth, Starlink V3 is expected to operate at 20kW, and they’re confident of going “right to 100 to 120.” On the skeptics, he channels Larry Ellison: “He’s out there landing rockets. I don’t see anybody else landing rockets.” The honest caveat survives: on repairing a broken rack in orbit, “until you have a floating Optimus, you don’t.”
  • The investable edge: inference goes orbital, training stays on Earth “for a long time,” and terrestrial data centers are “valuable for my lifetime” — America “is going to suck as hard as it can on every energy source,” and same for compute. But if you’re a power-and-cooling supplier massively ramping capacity that lands just as “all of the silly skeptics start to understand that orbital compute is very real — it’s worth thinking long and hard about that.”

4. Wafers: TSMC may single-handedly prevent the bubble — and TerraFab is coming

  • The wafer constraint is culturally load-bearing: TSMC’s leaders see themselves as inheritors of “Morris Chang’s sacred legacy” and the silicon shield. Twenty years ago at Science Park they told him catching Intel was “such a beautiful dream, but it’s a dream for our grandchildren” — and they did it. And Jensen has never had a contract with Taiwan Semi: “They do business on what seems fair in handshakes.”
  • Every foundational technology — canals, railroads, the internet — produced a bubble (Carlota Perez’s framework): markets correctly identify the technology, everyone turns bullish, supply overshoots, crash. There’s what he calls a breakdown in diversity. The fundamental differences this time: the build-out is “still overwhelmingly funded out of operating cash flows,” valuations are saner, and every GPU runs at 100% utilization versus 99% of fiber dark. But “based on the last 200 years… we should expect a bubble.”
  • The cautionary tale he keeps: George Vanderheiden, the great Fidelity PM who fought the ‘99 bubble — 40% tobacco, 40% homebuilders — retired in early 2000 because clients said “George, you’re out of step,” then outperformed the Nasdaq “by like 20 or 30x over the next 3 years.” Vanderheiden’s line: “Being early is the same thing as being wrong.”
  • The one thing to watch: TSMC’s capacity decisions. “If Taiwan Semi did what Jensen wanted, Nvidia could sell $2 trillion of GPUs in ‘26 or ‘27. Maybe 2½ trillion. Maybe 3” — enough to trigger overbuild. If no bubble comes, “we need to throw a party for them because they will have single-handedly prevented a bubble.” The risk: Intel or Samsung — “they’re not going to stay disciplined. They will break” — with TSMC’s leading-node edge at 9–15 months, there’s a Goldilocks zone between starving a second source and enabling one.

5. TerraFab: the A-teams follow Elon

  • The TerraFab — a SpaceX (and, he believes, Tesla) JV to build the world’s largest fab in America — he thinks will succeed, he argues, for three reasons: the Intel partnership delivers 50 years of institutional knowledge at “three to five quarters behind the front”; the semi-cap A-teams (ASML, KLA, Lam, Applied) will show up for Elon just as they once showed up in Taiwan — “they don’t like having a monopsony”; and it’s different enough not to alienate TSMC.
  • Talent is the third leg: Elon is “a living deity in China, Taiwan, South Korea, and Japan,” and Baker predicts a literal Taiwan town, Japan town and Korea town in Texas — favorite restaurants relocated, staff and all — because “the best engineers want to work for Elon, especially in hardware engineering. That’s just not the way the people who run Intel and Samsung think.”
  • On lead times, he shrugs: “Everybody else is taking 3 years to build a data center. He built one in 122 days.” Samsung had to give Elon an office in their Texas fab because he was so unhappy with their pace.

6. The three questions: frontier premium, the bitter lesson, continual learning

  • The surprise of the cycle: “an overwhelming amount” of model-layer economics accrue to frontier tokens — and you need a hypothesis on whether that persists. His own experience made him open-minded: Gemini 3.1 Pro was “mind-blowing” at launch and today is “Intolerable. Intolerable.” Companies prototype on frontier and sometimes deploy open source, but the frontier premium is a fact — for now.
  • The Pareto frontier (intelligence vs. cost) is his single most important lens on AI labs, and it flipped: Google dominated it 9 months ago; now Anthropic and OpenAI dominate, Grok 4.3 sits on it as “the best lowest-cost 500-billion-parameter model,” and Gemini 3.1 is “hanging on” — “if I were to bet, I’d bet they’re subsidizing that out of pride.” Root cause, from last episode: Google surrendered per-token cost leadership with conservative TPU v8 design choices while Nvidia stayed aggressive.
  • The biggest risk to the whole trade is a violation of Richard Sutton’s bitter lesson — human ingenuity beating brute compute. March’s scare, “turbo quad,” a year-old Google memory-optimization paper released mid-negotiation with Micron, Samsung and Hynix over a DRAM LTA (“what people do is always more important than what they say”), went viral as “DRAM is cooked” — yet he “was unable to find a single AI engineer on planet Earth who believed turbo quad would have any impact on DRAM demand.”
  • Where he breaks with the labs: builders are skeptical of bitter-lesson violations, but he thinks “we are very close to ASI. And who knows if the bitter lesson holds for 400-IQ models” — if you get to ASI, the first thing it probably wants is to be smarter and have more resources; it may make itself more efficient, and “the bitter lesson literally includes humans in it.” Question three is continual learning — today’s crude variant is just RL in mid-training on verifiable tasks; a human touches fire once, a model “needs to put its hand in the fire a million times” and then have designers build a fire into the next RL gym. Dynamically self-updating weights would mean “a really fast takeoff” — and people seem confident it’s “just around the corner.”

7. Pay by the drink: the pricing shift that funds well over $200 ARR

  • His telecom-analyst frame (‘05–‘07): cellular was a great growth industry under fixed-plus-usage pricing and stopped being a great growth industry when everyone went all-you-can-eat. “AI is just shifting from all you can eat to pay by the drink” — and people really like to use AI, “particularly now that one person can have 100 agents working.” This shift plus new compute is probably why OpenAI and Anthropic should exceed well over $200 in ARR this year.
  • Practical corollary for investors: the $250-a-month plan no longer shows you the frontier — “you’re getting severely rate limited. You’re getting a lobotomized version of the AI.” To understand what frontier AI can do, even for non-coding work, “you need Claude Code or Codex, and you need to be on an enterprise plan” — usage-based, so the model produces the tokens it actually thinks the answer needs. His conscience note: “It’s sad for the world… if you can’t afford that, you’re not at the frontier.”
  • Harnesses matter more than he once implied: “harness engineering is not as important as the model, but it really matters,” and harnesses and models are increasingly co-developed — context, memory, state, tools. “Even simple versions make an incredible difference.”

8. Chip startups: do something different AND hard — and save private credit doing it

  • Chip design has an iron triangle like tank design (attack, defense, mobility — the Merkava optimizes defense, Russian tanks mobility): trade-offs bounded by physics as embedded in TSMC’s design rules. TPU, Trainium and AMD are all “essentially trying to be a better GPU” — Trainium currently probably doing best (“they’re tugging on Superman’s cape”), though Trainium 3 must ramp because MoE inference economically requires its switch scale-up network; on AMD’s MI450, “we’ll see.”
  • His rule of thumb: 1% market share is worth $100 billion — a fine venture outcome — but Jensen’s implicit response is “if somebody does something different and it gets to 1 or 2 or 3% share, we’ll make that chip.” So the bar is different and hard. The disaggregation of prefill and decode opens the canvas — his colleague Andrew Fox’s image: “Prefill is loading the cannon, decode is firing it” — prefill is memory-capacity bound, decode memory-bandwidth bound. And never believe a startup’s “special TSMC process access”: “Jensen saw that process when it was a twinkle in Taiwan Semi’s eye.”
  • Cerebras is his exemplar of hard-and-different — wafer-scale computing, three generations of chips to get right, trying to address its shoreline-IO constraint with an optical wafer on top. “Everybody’s going to get funded after the Cerebras IPO.”
  • The sleeper macro call: disaggregation means GPUs get 10–15-year lives — put a Cerebras system or Groq LPUs (“that Nvidia acquired”) in front of a Hopper or even an Ampere for prefill and run it “until it melts.” That “may single-handedly save private credit,” which underwrote GPUs at 3–4 years, and cutting financing from CoreWeave’s low-7s toward 5–6% “mathematically changes the cost of financing this build-out.” Related, via his friend Jamin at Coatue: sellers of shortage are beating buyers — and since CPUs matter far more in an agentic world (orchestration, tool calls), the hyperscalers’ giant installed CPU fleets may claw back some of that.

9. Applications and open source: the token path and a new prisoner’s dilemma

  • At the application layer, “forget value accruing — trillions of dollars of value has been destroyed by AI,” even counting Cursor and Cognition. In Jensen’s five-layer cake, profits accrue to energy, data centers, chips and models — not applications. His venture filter: is the idea obvious to the world before you can build scale? Amazon’s e-commerce dominance killed every VC-backed lookalike; Wayfair survived because “they did something hard.” Founders in AI “are really struggling, man” — betting on niche data moats the model companies may simply run over.
  • Jamin Ball’s rule, which Baker endorses: “If you’re not in the token path and you’re not in some really niche thing, life may be hard.” The counterweight: coding was “righteous focus” — Cursor, Cognition and Anthropic focused when OpenAI did everything — and Replit’s founder’s framing stuck with him: coding may be “the shortest path to ASI,” because with code you can write yourself anything. If frontier-token returns ever compress, “there’s going to be an explosion in value creation at the application layer.”
  • Open-source frontier today is “Chinese models with stolen American tokens” — he’s told DeepSeek (one version) used only 150,000 reasoning traces; there are many ways a Chinese company could launder this across APIs; American labs are racing on anti-distillation tech. This, he argues, is partly why Mythos wasn’t broadly served: beyond the compute shortage, “they did not want it to be distilled” — better to distill it themselves and RL the next model.
  • The new prisoner’s dilemma: do you release your frontier model via API at all? If every lab abstains, Chinese open source catches up; if one defects, it takes the revenue, and “resources equal intelligence, so they’ll start to pull ahead” — same game theory as TSMC/Samsung/Intel. And Jensen “can probably get pretty close to the frontier whenever he wants” with his own model (likely Nemotron) — the logical counter-move that keeps open source a set distance behind. Open source isn’t free, either: tokens cost energy and the model companies “almost always get a revenue share.”

10. Diversity breakdown, broken baskets, and the mega-cap scorecard

  • His worry: a “breakdown in diversity” — “I don’t know of anyone like me who’s not really bullish on DRAM. No one.” Cross-sectionally, “the valuations flat out do not make sense”: semi-cap equipment at 40x next quarter’s annualized earnings vs. DRAM at mid-single-digits (last cycle’s peak gap ~5 vs. 12), and Nvidia — near its cheapest relative level in 10–12 years — is impossible to square with GE Vernova’s multiple. Meanwhile shortage economics mean the lowest-quality companies are performing best, exactly like high-cost commodity producers in a bull market, bid “to the moon” by retail accounts on X. Some of the ‘24–‘25 “nuclear bubble and quantum bubble” silliness has migrated into small caps. “I just wish there were more AI bears.”
  • Structurally, the AI trade stopped trading as one factor in January: scale-up networking ripped while scale-out fell, DRAM massively underperformed NAND and HDDs, “you couldn’t hedge your memory anymore with semi-cap.” His favorite hunting ground now is miscategorized names: Astera sat in “copper loser” baskets, but its biggest coming product is a switch — “definitionally, if you’re a switch company or an accelerator company, you cannot be a copper loser.”
  • The scorecard: Google lost the TPU edge but has the most compute, genuinely valuable YouTube data for robotics, and GCP “going crazy” — though if Google I/O (five days away) doesn’t at least slightly leapfrog OpenAI or Claude, “that’s interesting.” Zuckerberg gets “immense credit” as the only true internet giant to become AI-first internally, with MSL’s first model nearly on the Pareto frontier — and “rates of change matter more than level” on three-year horizons. Amazon is strong on Trainium plus real robotics P&L efficiency coming to retail within 18 months.
  • On Microsoft: Satya “did go from ‘we’re going to make Google dance’ to being the product manager of Copilot in like three years.” Microsoft “flinched for a moment in early ‘25” and lost allocations — but Baker calls the current move courageous: using compute for its own products and models rather than serving OpenAI, even though “Microsoft would probably be an $800 stock today” doing the opposite. His lingering counterfactual: during the coup attempt, does Satya wish he’d backed Ilya and Mira instead of Sam? Last observation: startup engagement is Amazon and Nvidia “by a mile,” then Google, Broadcom in its own ASIC lane — and “essentially zero” from AMD, Microsoft and Meta, which he thinks becomes a real disadvantage since “some of the best teams are no longer at big public companies.”

Coda: the event horizon

  • His preparation for Mythos 3 and 4 starts defensively: over-invest in cybersecurity, and “everybody needs to have a safe word” — one that can’t be socially engineered — against a FaceTime call that is “an utterly accurate simulation” of your child asking you to wire a million dollars. Darker still: with rising political violence in America (“it is terrible that someone threw Molotov cocktails at Sam Altman’s house”), he worries AI’s politicization puts its public figures at physical risk — “a higher variance, higher beta, higher risk world.”
  • The Last Samurai is his frame for his own career: the hero fights for the samurai and is massacred by “a peasant with a machine gun.” “The machine gun is here. And if we do not all become masters of the machine gun, we’re going to get mastered.” He’s optimistic a 50-year-old veteran has a long window of advantage wielding it — his most useful agent today, he admits to Patrick, is one that summarizes the 6 hours a day of podcasts in his job description, plus a first-pass proxy reader that flags PSU vs. RSU incentive changes.
  • Geopolitically: Ukraine “is really starting to win,” and not mainly on drones — “they have the best battlefield AI outside of probably America and Israel” — which is great for America but destabilizing for adversaries processing it; his hopeful read is a new Pax Americana, citing 1945: America had the bomb alone, could have controlled the world, and instead rebuilt Germany and Japan. He remains “an AI optimist maximalist” — the story of someone whose daughter’s rare mutation met an AI-agent-discovered on-market drug — “but it’s an event horizon… a discontinuity we need to navigate as society. It is a little dystopian that the best AI is only available to people with a lot of money. We need to solve that.”

Verification Notes

The raw captions leave the unit after “$200” unstated and render Krishna’s statistic only as “500% in DR”; the intended model name rendered as “Muse” is also unresolved.