Satya Nadella – How Microsoft thinks about AGI
Satya Nadella – How Microsoft thinks about AGI
Summary
- Microsoft deliberately walked away from being the biggest AI hoster. Dylan Patel’s numbers: Microsoft was on pace to pass Amazon by 2026-27 and hit 12–13GW by 2028; after the pause it’s ~9.5GW, and Oracle goes from 1/5th Microsoft’s size to bigger by end-2027 at 35% gross margins. Nadella’s defense: “We didn’t want to just be a hoster for one company” — the calculus is “not what you do in the next five years, but what you do for the next 50,” and fungibility across training/inference/geos beats gigawatts locked to one chip generation (“your entire network topology goes out of the window” on one MoE-like breakthrough).
- Coding AI is the tell on Microsoft’s competitive position: GitHub Copilot went from ~100% of a $500M category to sub-25% share of a $5–6B run-rate market in one year (Claude Code, Cursor ~$1B each; Codex $700–800M). Nadella embraces it — “there’s no birthright here… Thank God” the competitors aren’t Borland — and points to the cloud precedent: lower share of a vastly bigger market. His counter-product is Agent HQ/Mission Control, “the cable TV of all these AI agents,” packaging Codex, Claude, Grok and others into one GitHub subscription.
- The model-vs-scaffolding margin fight is the episode’s core disagreement. Nadella argues frontier labs face a “winner’s curse” — “one copy away from being commoditized” by open-source checkpoints, with data liquidity and scaffolding letting the app layer vertically integrate down into models. Dwarkesh’s pushback with numbers: Anthropic’s inference gross margins went from below 40% to north of 60% this year despite broad competition — the margin is expanding at the model layer, not the wrapper.
- The OpenAI deal terms include broad access: Microsoft has access to OpenAI’s models for seven more years and, on chips, “all of it” — every piece of OpenAI’s silicon and system-level IP except consumer hardware. OpenAI’s API (PaaS) is Azure-exclusive, including “stateful” partner deals — a Salesforce-style co-trained model deployed on AWS is not allowed, with only narrow carve-outs like USG.
- MAI is deliberately capital-efficient, not behind by accident: the text model debuted ~13 on LMArena trained on only ~15,000 H100s, image model #9; Nadella refuses to burn flops “duplicative” of the GPT family he already has access to, while still promising a “world-class superintelligence team.” Maia 200 “looks great,” but Google shipping 5–7M TPUs vs Microsoft’s small orders reflects his bar: the biggest competitor to any custom accelerator “is kind of even the previous generation of Nvidia.”
- Hyperscaling is now “a capital-intensive business and a knowledge-intensive business” — software is the moat on capex: 5x/10x/40x tokens-per-dollar-per-watt gains on a given GPT family, 90 days from data-center handover to live workload (Jensen’s advice: “speed-of-light execution”), and research compute should be accounted for “as R&D expense” with everything else demand-driven.
- Sovereignty and trust as the macro overlay: the US is 4% of population, 25% of GDP, 50% of market cap — a ratio built on trust that, if broken, “is not a good day for the United States.” Against Chinese competition, the winning feature “is not even the model capability, maybe. It is, can I trust you, the company… your country, and its institutions to be a long-term supplier.”
Deep dive
Not yet available upstream; scheduled sync will retry.