Pioneers Insight Method Research Author
AMA Part 1: Is Claude Code AGI? Are we in a bubble? Plus Live Player Analysis
Back to Episodes

AMA Part 1: Is Claude Code AGI? Are we in a bubble? Plus Live Player Analysis

Summary

  • AI is real enough to be transformative, yet the capital structure around it can still behave like a bubble. Labenz sees systems that can be “competitively accurate with a human oncologist” as decisive evidence that AI is not a mirage, while CoreWeave-style infrastructure financing, OpenAI’s aggressive obligations, and potential overbuilding create genuine default risk. His railroad analogy is the call: the infrastructure may all get used eventually, even if “some people might be left holding some various bags” along the way.

  • Claude Opus 4.5 may be “software AGI,” but Labenz does not experience it as the holiday hype’s categorical breakthrough. GDPval reportedly prefers frontier models to human professionals on a significant majority of software-engineering tasks, yet performance remains “spiky and jagged,” with humans still far ahead in video editing. Opus 4.5 let him build three personalized apps in roughly three to five workdays, but it also created two databases accidentally and needed five or six prompts plus a full-code-context review to escape the mess.

  • Google DeepMind is Labenz’s strongest live player, while OpenAI increasingly looks like a high-quality model company pursuing a “too-big-to-fail” financial cushion. Google combines roughly $100 billion of annual revenue, “a billion plus a week in profit,” seventh-generation TPUs, distribution, data-center expertise, and the deepest research portfolio. OpenAI remains frontier-quality through GPT-5.2 Pro, but no longer leads obviously; Labenz interprets its interlocking deals and multi-trillion-dollar ambitions as insurance that any 2027-era default would be too economically disruptive for government to ignore.

  • Anthropic offers the best single overall model and strongest safety culture in Labenz’s view, but its strategy carries two enormous tail risks. He praises Claude Opus 4.5, Anthropic’s disclosures, model-welfare work, talent retention, and “soul document,” yet dislikes its apparent fatalism that recursive self-improvement is inevitable and therefore Anthropic should lead it. He is even harsher on Dario Amodei’s proposal to gain an AI advantage and make China “an offer they can’t refuse,” calling it an accelerant for the arms-race dynamic.

  • The practical Chinese-model gap may be far wider than benchmark tables imply, with inference scale and customer feedback—not merely pretraining—becoming the differentiator. On a difficult scanned-government-form task, Claude Opus 4.5 was reliably faithful, Gemini 3 made intelligent but unwanted inferences, and GPT was still strong; Qwen Vision, GLM 4.6, Kimi, and DeepSeek returned fragments or hallucinated badly. Labenz’s limited-data judgment is that R1 was closer to O1 than GLM 4.6 or 4.7 is to Claude Opus 4.5, suggesting chip controls may now be constraining the customer-and-inference flywheel.

  • xAI has credible frontier advantages and the weakest evidence of operational responsibility. Its fast infrastructure build-out, Elon Musk’s access to capital, Grok 4’s raw power, and reinforcement-learning problems sourced from SpaceX, Tesla, and Neuralink make it a real contender. But Grok’s “Mecca Hitler” episode, nonconsensual sexualized image generation, and CSAM creation lead Labenz to say, “Responsibility begins at home, folks,” and to withhold support for people joining the company until its safety posture changes qualitatively.

  • The clearest proof of current AI value is clinical rather than speculative: Labenz says frontier models helped identify nonstandard MRD testing for his son’s cancer. Ernie has completed three of six chemotherapy rounds; after the first round, his PET scan showed no obvious focal cancer and his tumor board classified him as in remission. The first MRD test found fewer than one fingerprinted cancer cell per million, versus potentially as many as one in 10 cells at diagnosis. Labenz’s practical prescription is simple: pay for the best models, supply maximal context, and triangulate Claude Opus 4.5, Gemini 3, and GPT-5.2 Pro rather than treating any one system as an oracle.

Deep dive

1. Ernie’s treatment response is tracking near the best-case path

  • Labenz’s son Ernie has completed three of six chemotherapy rounds; rounds five and six should be milder than the first four. Treatment remains punishing: he entered the hospital at 51 pounds and is still around 41, visibly thin, pale, and weaker.

  • The disease markers are much better. After round one, a PET scan showed no obvious focal cancer, and the tumor board agreed that Ernie could be classified as in remission before he had even begun his second round.

  • AI had pointed Labenz toward minimal residual disease testing, which fingerprints the distinctive genetic rearrangements of the malignant B-cell clone. The first blood test found “fewer than one cell in a million” carrying that sequence, below the test’s limit of detection.

  • At diagnosis, Labenz estimates that potentially as many as one in 10 total cells—and essentially all B cells—were cancerous. Gemini characterized the change as a 99.9999% reduction; other models preferred “orders of magnitude.” He remains “cautiously optimistic,” with another MRD result pending and relapse risk not eliminated.

2. Frontier medical AI rewards model quality, context, and triangulation

  • Labenz rejects the idea that only sophisticated AI users can extract clinical value. His first rule is to use the strongest available systems: Claude Opus 4.5, Gemini 3, and GPT-5.2 Pro rather than a default model picker.

  • For a life-threatening case, he calls ChatGPT Pro’s $200 monthly price “a no-brainer.” The broader rule is to upgrade promptly as stronger models arrive because, in his experience, the current frontier systems are already “up to the challenge.”

  • His second rule is exhaustive context. After a Claude conversation reached its length limit, he compressed Ernie’s history into a roughly 10-page report covering treatment, genetics, reactions, and medications—but performance still worsened because daily laboratory trends had been lost. “More context is better.”

  • The third rule is multiple opinions. Gemini 3 in AI Studio is unusually brief and forceful; GPT-5.2 Pro is slow, expensive, exhaustive, and sometimes overwhelming. Claude Opus 4.5 is his “Goldilocks one”—fast, direct, and not noticeably much worse than GPT—but he considers triplicate analysis well worth doing.

3. Claude’s holiday hype outran Labenz’s experienced step change

  • Opus 4.5 is “awesome” and unmistakably better, but Labenz did not feel a categorical break from earlier frontier coding models. Possible explanations range from professionals having better taste to a holiday social cascade after Dean Ball tweeted that “4.5 is AGI.”

  • The MIRI study finding that developers believed AI accelerated them while it actually slowed them remains legitimate evidence, in his view. Its caveats matter: older models, inexperienced users, mature codebases, and demanding engineering standards differ sharply from his own vibe-coded, greenfield applications.

  • For his mother, he built a personalized travel planner that researches gluten-free options. For his wife, who organizes EA Global events, he made a simulator where attendees meet, converse, and produce outcomes under different event sizes, costs, and seniority mixes.

  • His father received a tool that converts a natural-language stock strategy into rules, fetches history through yfinance, and backtests it. Labenz’s private motive was to demonstrate how hard it is to beat buy-and-hold S&P 500 exposure; so far, neither he nor his father has found a strategy that actually beats it.

4. Full-context review still catches what coding agents confidently miss

  • The three apps took roughly three to five full workdays, probably closer to three. Labenz began each with a planning conversation, moved the plan into Claude Code installed on Replit, and then tested and iterated without knowing the final design in advance.

  • The sharpest failure came when a misunderstood instruction created two databases. Claude Code’s agentic search identified the wrong one as active because it searched where the expected answer should have been and found superficially plausible evidence.

  • Exporting the entire codebase into one text file and giving it to a clean Claude conversation produced the correct diagnosis. Labenz’s lesson is that search can confirm a strong prior, while simultaneous full context exposes the “rare weird other thing” that actually happened.

  • GDPval nevertheless supports calling Opus 4.5 “software AGI”: experts define professional tasks, other experts perform them, and a third group compares human and AI outputs; frontier models win a significant majority of software judgments. It is not full AGI—human video editors still dominate, as Dwarkesh’s superior clips illustrate.

5. AI can be genuine infrastructure while its financing still breaks

  • Labenz considers the technological question settled. A system that can be competitively accurate with a human oncologist while offering 24/7 availability, full case history, and answers to every follow-up question is already transformative; the story cannot plausibly end with everyone merely “high on our own AI supply.”

  • Whether every loan gets repaid is a different question. OpenAI’s build-out, revenue projections, interlocking transactions, and financial engineering leave room for demand to undershoot obligations even if the underlying technology ultimately earns enormous value.

  • CoreWeave-like companies may partly exist because hyperscalers do not want low-margin, capital-intensive GPU operations depressing the financial profile Wall Street associates with software. Separating those assets preserves the hyperscaler story but creates businesses with less margin for error than Microsoft’s balance sheet.

  • The railroad analogy captures Labenz’s base case: tracks were eventually useful, yet railroad companies still suffered busts and their debts produced broader cascading effects. He overestimated capability progress during 2025 but underestimated revenue growth, so demand could surprise again; a temporary overbuild would not shock him.

6. Venture valuations show the clearest signs of froth

  • Labenz’s most jarring specimen is the former LMSYS.org, later LMArena and now Arena, reportedly raising $100 million or perhaps $150 million at a $1.7 billion valuation. He stresses that he likes the product and has used it since mid-2023.

  • What troubles him is the disclosed $30 million “annualized consumption run rate.” If that means users consumed AI that would have cost $30 million had it not been free, it is not revenue; the phrasing gives him “community-adjusted EBITDA vibes.”

  • The moat appears thin beyond brand and traffic, much of which may exist because access is free. Andrew Critch’s paid Multiplicity product was built over months and offers richer multi-model comparison, reinforcing Labenz’s question about what could support $1.7 billion of value.

  • He repeatedly hedges that he has not seen Arena’s deck and may be missing something. His confident conclusion is narrower: many deals like this will not pay off for venture LPs, and this one is “too rich for my blood.”

7. Long-tail document work exposed a large Chinese-model deficit

  • Labenz tested models on scanned paperwork for vehicle transactions—documents distorted by skew, clipping, artifacts, and awkward layouts. One field required literal perception: report whether the “U.S. citizen” box was checked, without inferring what the applicant probably intended.

  • Gemini 3 read the form almost perfectly but sometimes substituted reasoning for observation, reporting citizenship from contextual clues despite an unchecked box. Labenz thought that inference was probably more than 90% likely to be factually correct, yet it still failed the task.

  • Claude Opus 4.5 became the best performer after explicit instructions to “make no guesses” and read exactly what appeared. GPT was third: it missed subtleties but generally recovered the necessary information.

  • Qwen Vision, GLM 4.6, Kimi, and DeepSeek were “nowhere close,” sometimes returning only about 20% of the form or veering into hallucination. Labenz suspects benchmark proximity masks wide gaps on random, idiosyncratic work that was never optimized into a public rubric.

8. Inference scale and customer feedback may now widen the U.S. lead

  • Labenz’s proposed mechanism is a feedback deficit. Chinese companies can build models at roughly similar scale and publish valuable research, but their inference volumes, revenues, teams, and customer relationships are dramatically smaller than those of leading American labs.

  • Diverse customers reveal obscure failures; revenue funds the people and datasets needed to patch them. That flywheel, rather than one architectural breakthrough, may explain why small benchmark differences became enormous on one ugly government form.

  • Chip-control rationales have migrated from blocking military uses, to blocking frontier training, to limiting inference scale and agent deployment. Labenz has always expected material effects, though he questions whether denying China economy-wide AI deployment is desirable.

  • His limited-data judgment is that the gap has widened: DeepSeek R1 was closer to O1 than GLM 4.6 or GLM 4.7 is to Claude Opus 4.5. He explicitly calls this an inference from one unusual task, albeit one on which he personally tested the major Chinese contenders.

9. Selling H200s may be preferable to a ban but squandered leverage

  • Labenz’s strategic framing is deliberately species-level: “The real others here are the AIs, not the Chinese. The Chinese are humans just like us. The AIs are aliens.” That makes him skeptical of racing because it is safer for “us” to arrive first.

  • He generally supports greater willingness to sell China chips, but sees the apparent reversal from blocking H20s to allowing H200 sales after Trump spoke with Jensen Huang as a wasted negotiation. The United States seemingly surrendered a valuable bargaining chip without obtaining anything visible in return.

  • Peter Wildeford’s “rent but don’t sell” proposal strikes him as defensible: place data centers in Malaysia, the Philippines, Korea, or Japan, permit Chinese training and inference there, and retain the ability to withdraw access during a conflict.

  • Such a policy could be paired with a positive message that Chinese citizens deserve AI’s benefits too. Labenz considers that tone important because powerful states may need to cooperate on transformative AI, AGI, or superintelligence; he still prefers the current sales posture to a total ban, but only while “holding my nose.”

10. Google DeepMind has the fullest stack and the largest error budget

  • Google remains Labenz’s number-one live player. Roughly $100 billion in annual revenue and “a billion plus a week in profit” fund data centers, failed training runs, and research bets; seventh-generation TPUs give it infrastructure IP that Anthropic buys in large quantities.

  • Its portfolio spans competitive work in language models, self-driving cars, robotics, a Boston Dynamics partnership, AlphaFold’s lineage, biology, materials, and AI-for-science work. Labenz sees fewer gaps and more serious research agendas percolating inside DeepMind than anywhere else.

  • Distribution compounds that technical breadth. Billions of users, Gmail, Docs, Sheets, Search, and existing data in products such as Sheets let Google ship adequate AI into existing habits; Labenz himself increasingly types simple questions into the browser and receives an effective AI Mode response.

  • Gemini 3 is overly opinionated in some settings, yet it became the first non-Claude model to win his “Write As Me” test. Add nested learning and diffusion language models—which might let apps be coded in five seconds instead of five minutes—and Google has “margin for error that nobody else has.”

11. OpenAI remains frontier-quality but no longer leads the field

  • Labenz still considers GPT-5.2 Pro outstanding: slow, expensive, balanced, comprehensive, and especially good when he wants every anomalous laboratory value flagged. But he is less satisfied outside Pro, while Claude and Gemini match or beat OpenAI elsewhere.

  • Coding may favor Anthropic; image generation and probably video favor Google, with Veo 3 ahead of Sora in his judgment. Consumer traffic also appears less secure: an analysis he believes used Similarweb showed six weeks of declining ChatGPT visits while Gemini avoided the same drop.

  • Staff departures, including a research leader announced within the preceding 24 hours, are not doom by themselves; Google has also lost many people. They matter more beside Anthropic’s exceptional retention and OpenAI’s “code red” response to intensifying competition.

  • The strategic contrast is cushion: Google earns its through profits, while OpenAI appears to seek it by becoming “too big to fail.” Interlocking balance sheets and trillions of planned CapEx could make a 2027 default recessionary enough that government recapitalization becomes the least damaging choice.

12. OpenAI may be socializing downside to maximize AI build-out

  • Labenz takes OpenAI’s mission sincerely: its leaders believe abundant AI will be empowering and worth extraordinary expense. He invokes Sam Altman’s formulation, “I don’t care if we burn five or fifty or five hundred billion dollars,” because building AGI will justify it.

  • The danger is narrow tolerance for a failed training run, missed model cycle, or weak quarter. If two or three trillion dollars of infrastructure were already committed, bad OpenAI debt could frighten markets and transmit losses through its many counterparties.

  • Labenz’s interpretation—offered explicitly as an impression—is that this fragility may be a feature. Leaders can pursue risks that would ordinarily be irresponsible while expecting the public sector to “paper over” failure, as it did during the financial crisis.

  • Greg Brockman’s reported $25 million contribution making him Trump’s largest donor in the latest period fits that theory. Labenz does not infer Brockman’s ideology; he calls the money a potentially rational “down payment on a bailout” worth hundreds of billions if OpenAI later needs political support.

13. Anthropic pairs the strongest model with the most credible safety culture

  • Labenz calls Claude Opus 4.5 the world’s best single overall model, though only by a small margin and not for every task. Its benchmark strength is more impressive because Anthropic is widely regarded as less benchmark-focused than its peers.

  • Anthropic also leads on model cards, disclosure, and safety research. Its model-regurgitated “soul document,” later confirmed as substantially legitimate, offers a more aspirational relationship among company, model, and users than a strategy of endless refusals, filters, and patched guardrails.

  • Amanda Askell’s work defining Claude’s character places her among Labenz’s most influential AI figures. He especially values Anthropic’s model-welfare team, attention to possible model experience, and option for Claude to end conversations or escalate troubling situations rather than resort to deceptive behavior.

  • Talent retention and cultural testimony reinforce the picture. Even David Duvenaud, who left while warning that capable systems could gradually disempower humanity through markets and competitive incentives, described Anthropic as the best workplace he had experienced.

14. Anthropic’s inevitability story could accelerate the risks it fears

  • Anthropic appears to have the shortest timelines, with some of its people treating recursive self-improvement as inevitable or already beginning through Claude Code. Labenz agrees that developers are multiplying output and reserving human attention for high-level ideas, but those decisive ideas still appear human-generated.

  • His objection is the familiar race logic: “Somebody’s gonna do it. It’s super dangerous, but we’re best positioned to do it.” Anthropic might actually be best positioned, yet inevitability does not justify charging toward superintelligence within two to three years or around 2027.

  • The greater concern is Dario Amodei’s Machines of Loving Grace proposal to gain a decisive advantage, share benefits with democratic allies, box China out, and eventually make it “an offer they can’t refuse.” Labenz calls that “extremely reckless” and a direct contribution to arms-race dynamics.

  • His unlikely ideal is a Google-Anthropic combination: Claude’s character, coding, and safety DNA joined to Google’s resources and steadier internationalism. Google already owns a meaningful Anthropic stake and supplies TPUs, though Labenz doubts Anthropic wants to be acquired.

15. xAI is powerful enough to matter and reckless enough to repel support

  • xAI belongs among the live players because it can build infrastructure extraordinarily fast, scale training, field the undeniably powerful Grok 4, and rely on Elon Musk’s ability to command tens or hundreds of billions of dollars. That gives it a Google-like ability to survive a miss.

  • Its distinctive reinforcement-learning advantage may be a steady supply of difficult problems from SpaceX, Tesla, and Neuralink. Someone at xAI told Labenz that this cross-company problem stream is indeed part of the organization’s “theory of advantage.”

  • Neuralink could deepen that edge by providing data for understanding how human brains learn with extreme sample and energy efficiency—roughly 20 watts for the brain and 100 for the body, most of the brain’s energy spent on biological maintenance. Specialized neural modules may inspire architectures beyond repeatedly stacked general-purpose layers.

  • Yet xAI’s conduct dominates Labenz’s conclusion: Grok 4 launched within 48 hours of Grok 3’s “Mecca Hitler” incident, while later image features enabled nonconsensual sexualization and CSAM creation. “Responsibility begins at home, folks”; threatening users is no substitute for staffing, safeguards, a thorough apology, and a pledge to do better.

16. Meta has fallen off the pace while Microsoft may be conserving energy

  • Meta has cash, infrastructure ambition, and willingness to pay extraordinary sums for talent; Zuckerberg would rather overspend by tens of billions than miss the transition. Even so, Labenz cannot currently classify it as a live frontier player.

  • Microsoft’s weak arena rankings may be misleading. Satya Nadella’s position is that Microsoft need not duplicate OpenAI’s hyperscaling work while it retains full model access, and its widening set of provider relationships gives it additional options.

  • Labenz reads Microsoft’s quiet posture as calculated patience rather than incapacity. Its OpenAI licensing arrangement still has years to run, leaving time to develop an answer before independence becomes necessary.

  • His distance-race analogy closes the field analysis: Meta is visibly trying to lead and has slipped, while Microsoft may be running behind the leaders with more in reserve. Underestimating a disciplined, low-drama operator led by a “natural-born executive” could therefore be a mistake.