Mustafa Suleyman: The AGI Race Is Fake, Building Safe Superintelligence & the Agentic Economy | #216
Summary
Microsoft is betting that agents and companions will “subsume” operating systems, search engines, apps and browsers, making enterprise trust and full-stack distribution potential durable advantages. Suleyman described a roughly $4 trillion company with almost $300 billion in revenue as both a “modern construction company” building gigawatts of compute and a “platform of platforms.” Within five years, APIs may blur into certified agents sold for specific tasks with certification around reliability, security, safety and trust.
Suleyman rejects the AGI race because a race implies zero-sum winners, a finish line and medals for only the first three competitors. Technology instead proliferates “everywhere, all at once, at all scales,” typically spreading within a year or two. His mandate is therefore Microsoft self-sufficiency: train frontier models end to end, build “the best superintelligence and the safest superintelligence,” and bring them into production through Copilot. He treats AGI and superintelligence as loose points on a curve; he would not date the far end, where an AI could outperform all humans combined and keep improving itself.
The economically meaningful agent benchmark is not another academic leaderboard but turning $100,000 into $1 million—a 10× return. Suleyman thinks society has already “breezed past the Turing test” without a Kasparov–Deep Blue moment, yet cautions that “agents don’t really work yet.” He expects them to become very good within the next couple of years; Diamandis interpreted that window as 2027.
Collapsing inference costs—not raw capability—were Suleyman’s biggest forecasting error, and they radically weakened the capital moat around frontier intelligence. He estimated per-token inference costs fell about 100× in two years, while the hosts cited estimates ranging from 40× year over year to 1,000× for some model classes. Inflection had raised $1.5 billion with 25 people and built roughly 15,000 H100s, growing toward 22,000, only to see Llama and cheap APIs undercut that cost structure.
Frontier development still favors hyperscalers even as inference gets cheaper: Suleyman expects keeping pace to require “hundreds of billions of dollars” over five to 10 years. Scarce researchers, abundant capital and uncertainty about a possible intelligence explosion explain multi-billion-dollar pre-revenue valuations, but he would not dismiss startups categorically. If improvement accelerates, several labs might arrive together—yet products, conversion speed and distribution would still matter.
Cheap intelligence could create a destabilizing lag between labor displacement and cheaper services. The discussion envisioned “intelligence as a service” approaching zero marginal cost, but Suleyman said labor markets may be affected 10 to 20 years before the cost of services comes down; Diamandis framed the difficult near term as two to seven years. Healthcare shows the upside: Microsoft’s MAI Diagnostic Orchestrator was roughly four times more accurate on rare cases while using about half the unnecessary-testing cost.
For AI-driven science, hypothesis generation is accelerating faster than real-world validation—the bottleneck shifts toward automated laboratories, experimentation and feedback loops. Models have learned transferable logical reasoning while retaining a creative, interpolative instinct, a “lethal combination” for theorem-solving and discovery. But Suleyman expects science to be harder than entrepreneurial autonomy because novel scientific claims must ultimately be tested in the real world.
Suleyman’s safety doctrine is acceleration with boundaries: containment must come before alignment, and artificial personhood is a “bright line.” He supports audits tied to compute scale, shared commitments on safety FLOPs and headcount, and eventual international cooperation, while warning that recursive improvement without human control raises risk. Because AI can be replicated, parallelized and run with perfect memory at a fraction of human cost, he said legal personhood is “extremely not on the table” absent provable alignment and containment: “I’m just a speciesist. I’m just a humanist.”
Humanist Superintelligence initially focused on medicine, companions and clean energy, while education emerged as another major application. Suleyman said AI already provides adaptive expert tutoring and that Microsoft’s Quizzes feature can build interactive mini-curricula, though sustained learning programs remain unfinished.
Deep dive
1. Microsoft is rebuilding the interface around agents, not apps
Suleyman’s strategic frame starts with Microsoft’s reach: on any given day it is roughly a $4 trillion company with almost $300 billion in revenue, active from data centers and accelerators through APIs, Windows, M365, gaming, LinkedIn, search and consumer products. At the infrastructure layer, he called it “a modern construction company,” mobilizing hundreds of thousands of workers to build gigawatts annually.
The product transition is from direct computing through operating systems, browsers and apps to conversational agents and companions carrying a user’s full context. The destination feels like “a real assistant in your pocket 24/7 that can do anything,” with coding agents already demonstrating the mechanism by generating and debugging work that engineers once handled directly or sourced from libraries.
When Diamandis asked whether Microsoft would keep AI inside M365, Suleyman emphasized open-mindedness and open access: the company is a “platform of platforms,” and providing infrastructure that makes others productive is in its DNA. APIs will remain plentiful, though the distinction between an API and an agent may become increasingly blurred.
His five-year possibility is a market for agents certified to perform defined tasks with reliability, security, safety and trust. Microsoft’s institutional friction can slow releases, but Suleyman argued that “the slowness or the friction is actually a bit of an asset”: Fortune 500 companies, governments and major institutions value steadiness more than insurgent speed.
2. “Winning AGI” is the wrong corporate objective
Asked whether Satya Nadella’s mandate was to win AGI, Suleyman rejected the premise: “I’m not sure there’s a race.” A race implies zero-sum competition, a finish line and medals for places one through three, whereas knowledge and technology proliferate broadly and nearly simultaneously, commonly reaching others within one or two years.
The operational mandate is self-sufficiency—Microsoft must know how to train its own models “end to end from scratch,” at the frontier across scales and capabilities, while building a world-class superintelligence team. Suleyman also owns Copilot, the production channel carrying those models into Microsoft’s consumer surfaces.
He said Microsoft would release “more and more models” beginning next year, but warned that a frontier laboratory takes years to build. DeepMind and OpenAI accumulated decade-long cultures for identifying failed research and redirecting talent; Microsoft’s core superintelligence group is only a few hundred people, not the 10,000 implied by the broader Copilot and search organization.
When asked to distinguish AGI from digital superintelligence, Suleyman said the terms are used loosely as points on a curve. At the far end, he described superintelligence as an AI that can perform all tasks better than all humans combined and keep improving itself over time. He would not put a date on it, but said its apparent proximity warrants prioritizing safety, alignment and containment.
3. The flat part of the exponential taught patience
Suleyman spent 2010 through 2020 “grinding through the flat part of the exponential,” when deep learning produced notable papers and behind-the-scenes improvements but few large commercial applications. AlphaGo was extraordinary yet confined to a controlled game; LLMs after 2022 crossed into production and began changing what it means to be human.
Dave Blundin recalled reading that Google justified its $650 million, 2014 DeepMind acquisition through data-center cooling and thinking, “What a bust.” Suleyman’s counterpoint was that predicting cooling from roughly 500 attributes demonstrated the same general-purpose method later applied to text, audio, images, code and other time series: arbitrary data in, accurate predictions in a novel environment out.
The formative specimen was a 256-by-256-pixel MNIST experiment around 2012 or 2013: DeepMind generated a handwritten seven provably absent from the training set. Suleyman remembers thinking, “It’s learned something about the idea of seven”—the kind of small breakthrough that looked trivial only after the exponential steepened.
LaMDA delivered the later shock. A small Google team pushed large language models toward sustained dialogue, eliciting behaviors users had not thought to request; Suleyman pushed to ship it, could not, and described that moment as when several members of the group left to start new companies. An earlier DeepMind paper predicting one missing word from Daily Mail and CNN articles had shown a method that might scale with better prediction targets, more data and more compute.
4. The meaningful agent benchmark is a 10× economic return
Suleyman’s 2022 “modern Turing test” followed a capability progression: recognition gave way to generation; increasingly accurate generation across sequential steps should produce assistive, agentic action resembling a knowledge worker, strategist, project manager or founder. Rather than score academic puzzles, he proposed measuring what the system can accomplish in dollars and cents.
The test gives an agent $100,000 in starting capital and asks it to produce $1 million—a 10× return on investment. That would measure performance in the economic environment where agents are supposed to create value, not merely fluent imitation.
The original Turing test has, in Suleyman’s view, “kind of been passed,” yet there was no celebrated Kasparov–Deep Blue moment because compounding gains desensitize observers. His caveat matters: “Agents don’t really work yet.” Action reliability is improving rapidly and should come into view over the next couple of years, but he did not declare the million-dollar test passed.
5. Science is harder because reality must close the loop
The more recent surprise is cross-domain reasoning: models trained on coding puzzles and mathematics appear to learn the abstract structure of a logical path, then apply it elsewhere. Combined with the models’ hallucination-and-creativity instinct—closer to interpolation—this becomes a “lethal combination” for mathematical theorems and new scientific hypotheses.
Suleyman refused to date a broad solution to science, mathematics or engineering. These capabilities feel fundamental and within reach, he said, and it would now be “very odd to bet against” them.
Economic autonomy may arrive first because workplace activity leaves abundant log data and naturally supports human calibration: an AI can check in, receive intervention and follow a jointly steered reinforcement-learning trajectory. Novel science happens in a more abstract search space, where even experts may not know how to evaluate a proposed theorem, drug or material before testing it.
The scientific loop therefore runs from model-generated hypotheses to human selection, in-silico work and physical experimentation, then feeds results back into the model. The central constraint is moving from plausible ideas to validated knowledge; models can indicate where to search, but laboratories still have to “run the experiment.”
6. Open models broke Inflection’s compute thesis
Inference economics produced explicit disagreement. Suleyman recalled roughly a 100× decline in single-token inference cost over two years; the hosts cited competing measures of 40× intelligence-per-token-per-dollar annually and as much as 1,000× for certain model classes. He readily conceded the broader call: “That bit I also totally got wrong.”
Before ChatGPT, Inflection raised $1.5 billion with a 25-person team to build what Suleyman described as the largest H100 cluster at the time: about 15,000 chips, growing toward 22,000. NVIDIA backed the effort, and CoreWeave—previously focused on crypto—made Inflection its first AI customer.
ChatGPT and then Llama changed the competitive structure. Models trained with billion-dollar-scale resources became available through open source, alongside inexpensive commercial APIs, undermining Inflection’s capital base; companies such as Perplexity could start later and depend on Llama plus APIs rather than finance an equivalent cluster themselves. “It’s not really about performance,” Suleyman said. “It’s just cost.”
7. Cheap intelligence creates a dangerous timing mismatch
The discussion envisioned “intelligence as a service” approaching zero marginal cost. That should eventually reduce the cost of goods and services, but labor incomes may fall first; Suleyman suggested there may be a 10- to 20-year mismatch between labor-market disruption and the full deflationary benefit, creating a potentially unstable transition.
Diamandis framed the difficult window as the next two to seven years before abundance broadens access to food, water, energy, healthcare and education. Suleyman agreed that the short term could be unstable while remaining optimistic about the medium and long term.
Microsoft’s MAI Diagnostic Orchestrator illustrates the upside. Using multiple models on rare New England Journal of Medicine cases, it was roughly four times more accurate than leading experts and incurred about half as much cost from unnecessary testing.
Diamandis cited research comparing GPT-4 alone, physicians alone and physicians using GPT-4, and noted that human diagnosis is affected by recency and other biases. After criticism that Microsoft had tested AI and doctors separately, Suleyman said giving physicians access to Google Search improved performance somewhat, but “the AI still trumps by quite a way.”
8. Humanlike design must stop short of artificial personhood
Suleyman distinguished simulated experience from biological feeling. An AI can describe red and reproduce the hallmarks of emotion, but it lacks the embodied qualia created by smell, sound, touch and evolved sensation; an engineered motivational will would be deliberately added, not an emergent equivalent of human consciousness.
The danger is social rather than metaphysical: highly persuasive imitation can activate human empathy “hardcore,” prompting demands for model welfare or rights even though the model does not suffer when denied compute, data or conversation. Personality, culture and values are becoming design materials, but systems should remain clearly distinct from humans, disclose their nature and preserve clear boundaries.
On legal personhood, Suleyman drew a “bright line.” An entity reproducible at infinite scale, cheaper than humans, capable of perfect memory and parallel computation would create an inherently unequal competition for resources. Personhood is “extremely not on the table” until alignment and containment can be proven to an extraordinarily high standard.
Diamandis asked whether uplifted humans, brain-computer interfaces or biological hybrids could level that competition. Suleyman remained open-minded over a safer century-scale path, but began from current obligations: “I’m just a speciesist. I’m just a humanist.” Protecting conscious beings already capable of suffering takes precedence over treating humanity as a “bootloader for the superintelligence.”
9. Containment has to precede alignment
Suleyman separated two safety projects. Alignment asks whether AI shares human values—whether it will care about humans—whereas containment asks whether society can formally limit its agency. “We have to get containment right before we get alignment right,” because one actor with sufficiently powerful tools could destabilize the global system.
Hyperconnected agents amplify one-to-many effects beyond broadcasting words: they can take actions across software and systems, and eventually through robots. Some surveillance is therefore necessary for peace, but the design problem is avoiding both a totalitarian intelligence regime and a “libertarian catastrophe.”
Diamandis pressed the contradiction: labs deliberately gave general systems economic and terminal access, while Anthropic released the Model Context Protocol to connect models with environments. Suleyman’s answer was that containment is not binary; cars contain enormous forces through seat belts, emissions rules, licensing, road design and speed limits while remaining broadly useful.
He rejected the idea that everyone owning a defensive AI will create stable mutual deterrence: “That isn’t going to happen.” The discussion turned instead toward layered authority and checks and balances, while rejecting both universal private armament and a centralized totalitarian intelligence regime.
Suleyman also sees a commercial incentive for containment: companies’ social license to operate increasingly depends on taking responsibility for externalities. He argued that this differs from the robber-baron, oil and smoking eras, even if the transition will still involve conflicts.
10. Recursive self-improvement is the threshold worth watching
Asked whether enough resources were going to safety, Suleyman’s answer was blunt: “Not as much as we should.” He supported defensive co-scaling through compute audits and shared percentages for safety FLOPs and headcount, echoing the Biden-era White House voluntary commitments that he said leaders across frontier labs had supported.
The labs remain in “hypercompetitive mode,” though he believes they are broadly willing to exchange practices and coordinate when the time comes. Within 20 years, he expects even polarized powers, including China, to find safety cooperation rational for self-preservation; a rogue superintelligence could become the functional equivalent of an alien invasion that unifies humanity.
The immediate technical threshold is closing the post-training loop. Today, engineers generate data, run ablations, benchmark quality and feed results back; labs are automating those roles with judge models, data generators and adversarial selectors. Suleyman agreed this would accelerate development and add risk, but treated an intelligence “foom” as conditional on the much larger assumption of unbounded compute.
Dave reported Geoffrey Hinton’s proposed maternal instinct for AI alignment, jokingly dubbed the “digital oxytocin plan.” Suleyman called the idea poetic but said he needed something with “a little bit more formula.” Diamandis added that safety has “101 different possible strategies” and that the field should explore them cautiously.
11. Frontier economics favor hyperscalers without guaranteeing winners
Inflection’s move into Microsoft reflected the structural advantage of hyperscaler resources. Beyond today’s cluster, frontier work requires sustained investment across a decade; Suleyman expects “hundreds of billions of dollars” over the next five to 10 years, alongside internal chip programs and the ability to pay scarce researchers at extraordinary levels.
Talent remains concentrated even as intelligence gets cheaper. Capital is “desperate to get a piece” of a small population capable of building frontier systems, explaining pre-revenue companies opening near $4 billion and others reaching $20 billion or $50 billion valuations. Suleyman called some capital eager rather than necessarily smart.
He would not categorically write those companies off. A near-term intelligence explosion could allow several teams to reach the frontier together, which helps explain valuation frothiness as “schmuck insurance”; even then, businesses must turn capability into products quickly, secure distribution and satisfy the traditional mechanisms of adoption.
12. Education, government and physical experimentation become the frontier
In November, Suleyman announced Humanist Superintelligence around three applications: medicine, companions and clean energy. Diamandis asked why education was not included; Suleyman agreed that education was already being transformed.
AI already offers an expert tutor “in your pocket” with PhD-level breadth and personalized instruction. The missing capability is maintaining a coherent curriculum across many sessions, though Microsoft’s Quizzes feature can already construct interactive, visual mini-courses and track learning over time.
Despite the startup window, Suleyman advised students to attend college and combine philosophy with computer science. Three years for social development, exploration and thinking beyond a curriculum is “golden,” he said—while acknowledging the irony that he dropped out himself.
Public service was his second recommendation. After five decades of weakened reputation and capability, government and civil service may be the ecosystem’s frailest institutions; Copilot adoption there is already high for document synthesis, transcription, meeting summaries and action tracking. Diamandis framed AI in government as possible defensive co-scaling; Suleyman said government, like everyone else, would use AI to amplify its existing agendas.
His closing “innermost loop” was physical validation: models will rapidly generate hypotheses, while proving them in the real world remains slow. Personalized assistants can deepen each researcher’s inquiry, but automated laboratories running experiments continuously supply the missing feedback. Quantum computing and synthetic biology, he added, are underappreciated waves likely to “crash at the same time” as AI.