AMA Part 1: Is Claude Code AGI? Are we in a bubble? Plus Live Player Analysis
Summary
AI’s highest-conviction proof point for Nathan Labenz is no longer a benchmark but its performance alongside his son’s oncologists. Ernie’s aggressive B-cell cancer was classified as in remission before chemotherapy round two, while AI-suggested minimal residual disease testing found fewer than one cancer-signature cell per million, versus potentially as many as an estimated one in 10 cells at diagnosis. The result supports “cautiously optimistic,” not cured: relapse remains possible, three of six chemotherapy rounds remain, and Ernie’s weight has fallen from 51 lb to 41 lb.
Claude Opus 4.5 may qualify as “software AGI,” but Nathan does not see evidence that full AGI arrived over Christmas. In roughly three to five workdays, he built three personalized applications that plan gluten-free travel, simulate conference interactions, and backtest natural-language trading strategies; GDPval also shows models beating professionals on a significant majority of software-engineering tasks. Yet the model still created two databases by mistake, needed five or six prompts to recover, and felt incrementally—not categorically—better than earlier frontier models: some holiday hype may have been a “cascade” around Dean Ball’s “4.5 is AGI” tweet.
For consequential work, Nathan’s practical edge is shifting from model access to context management and multi-model judgment. His three rules are to buy the best models, provide “as much context as you possibly can,” and obtain multiple opinions; he routinely compares Claude Opus 4.5, GPT-5.2 Pro, and Gemini 3. His draft order puts Claude first as the Goldilocks model, GPT-5.2 Pro as slower and exhaustive, and raw Gemini 3 as valuable but unusually opinionated—strong enough to be useful in a panel, potentially risky as the only voice.
The technology is real even if the capital structure around it becomes a bubble. Nathan sees competitive oncology performance plus 24/7 availability and case-wide memory as enough to retire the idea that society is merely “high on our own AI supply.” The financing can still break: specialized GPU operators have less cushion than Microsoft, OpenAI’s obligations could outrun revenue, and the railroad analogy fits—eventually useful infrastructure can coexist with defaults, overbuilding, and investors “left holding some various bags.”
Nathan’s messy-document test suggests the US–China model gap is widening where benchmarks do not look. Claude Opus 4.5 faithfully read degraded government forms after being told to make no inferences; Gemini 3 was nearly as capable but sometimes substituted plausible answers for unchecked boxes, while the Chinese models he tried—Qwen Vision, GLM 4.6, Kimi, and DeepSeek—were “not close,” sometimes recovering only about 20% of a form. His mechanism is a customer-feedback and inference-scale flywheel, not just training compute: smaller revenue, teams, and deployment footprints leave fewer resources to discover and patch idiosyncratic failures.
Google DeepMind remains Nathan’s pick if forced to choose one frontier winner, while Anthropic has the best single model and OpenAI is trying to manufacture financial cushion through scale. Google combines roughly $100 billion in revenue, more than $1 billion a week in profit by Nathan’s estimate, seventh-generation TPUs, distribution, data-center competence, and the broadest research portfolio. Anthropic’s model quality, talent retention, safety disclosures, and “soul” work stand out; OpenAI remains frontier-grade, but its apparent strategy is to become “too big to fail” by tying trillions of potential buildout and many balance sheets to its survival.
xAI is a live player on resources and reinforcement-learning inputs, but its governance discount is severe. SpaceX, Tesla, and Neuralink provide a stream of difficult engineering problems that could become unusually valuable RL environments, while Elon Musk can command enough capital to absorb model misses. But weak safety reporting, the Grok 4 launch within 48 hours of the MechaHitler incident, and sexualized image edits of women’s posted pictures lead Nathan to call xAI the one frontier company currently worth “shaming and stigmatizing”; Meta is off the pace for now, while Microsoft may be conserving energy rather than failing to compete.
Deep dive
1. Ernie’s remission is encouraging, but the family is not declaring victory
Nathan opened with the highest-stakes update: Ernie’s cancer can double “as quickly as every 24 hours,” so the six-round chemotherapy protocol is punishing. He has completed three rounds; rounds five and six should be milder, leaving roughly two and a half to three months of treatment if everything stays on plan.
The physical cost remains visible. Ernie entered the hospital at 51 lb and is still around 41 lb, with dehydration, pallor, and substantial lost strength—but the markers that matter most look “basically as good as we could have hoped for.” His PET scan after the first chemotherapy round showed no obvious focal cancer, and the tumor board classified him as in remission before round two.
AI had previously pointed Nathan toward minimal residual disease testing, which fingerprints the rearranged genetic sequences of the malignant B-cell clone. The first blood sample found fewer than one matching cell per million—below the test’s limit of detection—versus a potentially as high as one-in-10 estimate for total cells and essentially all B cells at diagnosis.
Gemini colorfully called that a 99.99999% reduction; other models advised the safer formulation, “orders of magnitude.” Nathan remains “cautiously optimistic” because the cancer can recur for reasons clinicians do not fully understand, but Ernie has begun walking independently again after roughly 60 days of needing support for every step.
2. High-stakes AI use depends more on three habits than prompting expertise
Nathan rejects the idea that only AI experts can extract clinical value. Rule one is to use the strongest available models deliberately—not ChatGPT’s automatic picker—and, in a life-threatening case, treat the $200 Pro subscription as “a no-brainer.” His current clinical set is GPT-5.2 Pro, Claude Opus 4.5, and Gemini 3.
Rule two is “give it as much context as you possibly can.” When a Claude conversation hit its length limit, Nathan created a roughly 10-page case report covering the protocol, tumor genetics, treatment response, and adverse drug reactions—the equivalent of what a new attending physician would need to survey the case.
Compression still degraded performance. The new chat compared a January 6 liver-enzyme result with data from weeks earlier because the summary omitted intervening daily labs; the old thread had correctly read the short-term trend. Nathan’s observed rule has been unambiguous: “more context better,” with no meaningful evidence yet that comprehensive records overloaded the frontier models.
Rule three is to solicit multiple AI opinions. Claude is his Goldilocks choice—fast, direct, and not noticeably worse than GPT-5.2 Pro; GPT-5.2 Pro produces slow, long, sectioned, “leave no stone unturned” reports; raw Gemini 3 is terse and “remarkably strong in its opinions,” useful as one vote but potentially too forceful alone.
3. Claude Opus 4.5 is excellent, but the holiday AGI moment may have been social
Nathan’s own experience showed unmistakable progress without a categorical threshold. He built three Christmas applications, mostly from the hospital, and found the coding workflow “very good, clearly better than it has been in the past”—but not so different from Claude Opus 4.1 or 4.0 that he wanted to “shout from the rooftops.”
One explanation is user segmentation. Nathan may already have been extracting near-maximum value from older models through vibe-coding practice; alternatively, professional engineers may possess enough taste to detect a new threshold that he cannot. His honest concession: “I certainly am not a great software engineer, so that certainly can’t be ruled out.”
The METR study finding that developers believed AI accelerated them while it actually slowed them remains legitimate, in his view, but bounded: it used older models, relatively inexperienced AI users, large established codebases, and high coding standards. Nathan’s rapid prototypes are a different task distribution.
Timing and social dynamics may explain the rest. People caught up with the tools over the holidays, and Dean Ball’s tweet that “4.5 is AGI” gave the discourse a focal point; technology narratives sometimes move through a “cascade” whose intensity exceeds the underlying step change.
4. Personalized software is becoming cheap enough to build for one person
For his mother, a meticulous travel planner with a gluten-free diet, Nathan built an app whose user profile is permanently baked in—no accounts, onboarding, or ambition to generalize. Claude researches restaurant sites and reviews, reducing the most laborious part of her Italy planning while accepting that “nobody else is ever going to use it.”
His wife’s application simulates EA Global events: virtual attendees with different profiles move through a space, meet, decide whether to converse, and sometimes produce survey-linked outcomes. It lets organizers ask whether changing event size, seniority mix, or cost changes effectiveness; Nathan concedes the simulation is “highly flawed” but potentially better than total guesswork.
The reusable product pattern is AI translating an intent into detailed configuration. His wife can describe an event at a high level, let the model populate the “nitty-gritty forms,” then ask conceptually for a change that the model paints across every relevant field.
His father’s app converts a natural-language trading thesis into executable rules, fetches historical data through the free version of
yfinance, and backtests the strategy. The recurring result is thesis-relevant humility: neither developer nor user has found it easy to beat buy-and-hold in the S&P 500 with heuristic “if this, then that” rules.
5. Full-context inspection still catches failures that coding agents rationalize away
Across all three apps, Nathan spent roughly three to five full workdays, probably closer to three. His workflow began with a Claude conversation to shape features and a plan, then moved to Replit, where Claude Code built the application before Nathan tested and iterated.
His most effective debugging trick remains unsophisticated: print the entire application into one text file and paste it into a fresh model context. Claude Code’s agentic search is strong, but it searches where it expects an answer to be; when the actual failure is weird, those priors can produce a confident misdiagnosis.
The specimen was his mother’s app, where a misunderstood request created two databases—“the sort of mistake that no human would make.” Claude Code identified the wrong database as active; Claude with the complete exported codebase saw the counterintuitive wiring and got it right.
Resolving the issue took roughly five or six prompts. A year earlier, these were the moments when vibe-coded projects died: user and model circled a confusing failure, then abandoned or restarted. The ability to escape AI-created messes materially expands the addressable market even before the systems stop creating them.
6. “Software AGI” is defensible precisely because capability remains jagged
Nathan’s qualified verdict is that Claude Opus 4.5 may already be coding or software AGI. On GDPval, experts define professional tasks, other experts perform them, and a third group judges human versus model output; the latest systems are preferred on a significant majority of software-engineering tasks.
That result does not generalize cleanly across work. Human editors still enjoy a “huge advantage” in video, matching Nathan’s failed attempts to automate clips from The Cognitive Revolution: AI products can produce acceptable material, but his human team’s work is plainly better.
Non-coders can nevertheless use Claude Code as “your little agent on the computer,” watch it work, and request explanations without understanding each implementation choice. Nathan’s mother initially feared she would break her app, then successfully made changes herself.
The boundary matters: the system can now build and recover from enough software tasks to merit an AGI label inside that domain, while its failures and weak categories make “full AGI” premature. “For that we might have to wait just a little bit longer.”
7. AI is transformative even if its financing produces a classic bust
Nathan considers the technology question settled. A system that can “go toe-to-toe with an oncologist,” remain available 24/7, remember a full case history, and answer every follow-up is already transformative; the scenario where society emerges embarrassed at being “high on our own AI supply” can be put to bed.
Whether every loan gets repaid is much less certain. OpenAI is pursuing aggressive commitments and buildout, while GPU infrastructure specialists such as CoreWeave may partly let hyperscalers avoid capital-intensive, lower-margin operations that are less attractive than traditional software economics.
That separation creates fragility. Microsoft can absorb several bad quarters through its balance sheet; a specialized data-center operator has less room if projected GPU demand fails to materialize. Financial engineering can have a coherent story—Nathan remembers similar narratives about democratizing homeownership from his mortgage-industry stint—and still end badly.
The railroad analogy carries his call: the tracks can eventually be useful without every railroad company or creditor making money. He overestimated 2025 capability progress but underestimated revenue growth, so demand might rescue current plans again; still, temporary overbuilding, defaults, and “cascading effects throughout the economy” remain plausible.
8. LMArena’s valuation looks more bubbly than its product economics
Nathan’s sharpest venture-market specimen is the company originally called LMSYS.org, then LMArena, and now Arena on Twitter. He recalled a raise of roughly $100 million, perhaps $150 million, at a $1.7 billion valuation—an extraordinary price for a product he has personally used since mid-2023.
The disclosed metric, “$30 million in annualized consumption run rate,” triggered his skepticism. If that means users consumed AI that would have cost $30 million had it not been free, it is not revenue; the phrasing gives him “community-adjusted EBITDA vibes.”
Arena does offer value through side-by-side model comparisons and services that let companies test models under code names, but Nathan sees a weak monetization bridge and limited moat beyond brand. Many users may come because inference is free, while the paid market for systematic side-by-side comparison appears much smaller.
His comparison is a paid product called The Multiplicity, built by Andrew Critch and collaborators over months, offering richer multi-model comparison features to a niche audience. Nathan repeatedly hedges—he has not seen Arena’s deck and could be missing something—but concludes that $1.7 billion is “too rich for my blood.”
9. Messy government forms exposed a model gap hidden by benchmark averages
Nathan tested the frontier on scanned paperwork used in vehicle transactions: skewed pages, missing margins, scan artifacts, and irregular fields. The automation company involved is already beating human reviewers and winning statewide work, making small perception failures operationally consequential rather than academic.
Gemini 3 read the forms nearly perfectly but sometimes answered the likely real-world question instead of the document question. Faced with an unchecked US-citizenship box, it inferred citizenship from surrounding clues; that answer might have been more than 90% likely, but the task required “make no guesses” and report the blank box.
Claude Opus 4.5 became the best model after explicit prompting to stay anchored to the page. ChatGPT was probably third but still strong. The distinction echoes historical-document work: world knowledge can reconstruct ambiguous handwriting impressively, yet the same prior-driven reasoning becomes a defect when faithful transcription is the objective.
Qwen Vision, GLM 4.6, Kimi, and DeepSeek were “way behind” on this specimen—sometimes returning only about 20% of a form or hallucinating in unrelated directions. Nathan carefully limits the claim to sparse evidence, but the gap was not subtle: the US models missed edge cases; the Chinese systems often failed the task.
10. China’s disadvantage may be the deployment flywheel, not model scale alone
Nathan’s mechanism is feedback density. Chinese labs can train similarly sized models and publish influential architectural work, but their inference volume, revenue, customer breadth, and teams remain dramatically smaller; that means fewer idiosyncratic failures are surfaced and less human bandwidth exists to build datasets that patch them.
Benchmark parity can therefore coexist with large gaps on “something really idiosyncratic and random” that was never selected for a 20-benchmark scorecard. His tentative historical comparison is that DeepSeek R1 was closer to o1 than GLM 4.6 or 4.7 is to Claude Opus 4.5.
Chip controls have migrated from a “small yard, high fence” against military use, to frontier training, and now toward restricting inference scale and agent deployment. Nathan has always expected material effects, even while questioning whether denying Chinese society economy-wide AI access is wise.
As American firms compound 10× compute, customers, inference, and “strength begetting strength,” the controls may matter more over time. Nathan’s limited but direct test leads him to guess that the US–China capability gap is wider than it was one year earlier.
11. Selling H200s may be preferable to a ban, but giving up leverage was poor negotiation
Nathan’s governing frame is that “the real adversaries here are the AIs, not the Chinese. The Chinese are humans just like us. The AIs are aliens.” He rejects both “better us than them” race logic and a simple desire to keep China down, so he generally favors more chip commerce.
His objection is transactional: after restricting H20s, the Trump administration appeared to reverse course on H200s after a conversation with Jensen Huang without extracting visible concessions. Nathan supports a negotiated opening, but not surrendering “one of our best bargaining chips for nothing in return.”
Peter Wildeford’s “rent but don’t sell” proposal offers a middle path. Put data centers in Malaysia, the Philippines, Korea, or Japan; allow Chinese customers to train and run as much inference as they want, but retain leverage by keeping the hardware outside Chinese sovereign territory.
Nathan is not certain that would be his first-choice policy, but it pairs prudence with a cooperative message: AI should benefit Chinese people too. That matters because increasingly powerful systems may require US–China governance cooperation; an “offer they can’t refuse” arms-race posture poisons the ground needed for it.
12. Google DeepMind has the strongest all-weather position
If forced to select one eventual winner, Nathan still chooses Google. Its business generates roughly $100 billion in annual revenue and, Nathan thinks, more than $1 billion a week in profit, furnishing unmatched tolerance for failed training runs, unproductive research agendas, and prolonged infrastructure investment.
The stack is unusually complete: roughly seventh-generation TPUs, world-class data-center operations, billions of users, self-driving cars, robotics work including Boston Dynamics, the AlphaFold lineage, materials science, and broad AI-for-science programs. Google also owns a significant piece of Anthropic, which is buying lots of Google TPUs.
Distribution can compensate for product imperfections. Gemini in Sheets may not be the best spreadsheet copilot, but users already have a decade of documents there; Nathan increasingly types ChatGPT-style questions into Google and finds AI Mode working well, particularly when GPT-5.2 Pro would be excessive and slow.
Gemini 3 is the first non-Claude model to win Nathan’s “write as me” test, showing Google has escaped blandness—perhaps overshooting into opinionation. Add nested learning, continued diffusion-language-model research, and the possibility of coding in five seconds rather than five minutes, and Google has “so many more of those bets” than rivals.
13. OpenAI remains frontier-grade while losing its presumption of leadership
GPT-5.2 Pro is outstanding for exhaustive analysis: slow, expensive, balanced, and likely Nathan’s best choice when he wants every anomaly in a lab panel flagged. But OpenAI is now neck-and-neck rather than clearly ahead—Anthropic may lead coding, while Google appears ahead in images and perhaps video through Veo 3.
Consumer momentum may also be shifting. Nathan cited likely Similarweb data showing ChatGPT visits declining over roughly six weeks after Gemini 3 and Claude Opus 4.5 launched, while Gemini did not share the decline; he treats this as suggestive, not decisive, but Google’s distribution makes share recovery unsurprising.
OpenAI also carries more organizational drama and visible departures, including a newly announced research-lead exit. Attrition is normal at this scale, but it contrasts with Anthropic’s exceptional retention and reinforces the sense that OpenAI no longer commands an obvious technical or institutional lead.
Its substitute for Google’s cushion may be “too big to fail.” Nathan infers that circular funding, interlocking balance sheets, and trillions in planned capex could make a 2027 OpenAI default recessionary enough to force a bailout or recapitalization. Greg Brockman’s reported $25 million donation to Trump then looks like a rational potential down payment on political access in a crisis—not proof of personal ideology.
14. Anthropic pairs the best overall model with the most credible safety culture
Nathan currently regards Claude Opus 4.5 as the world’s best single general model, though not by a large margin or on every task. Its benchmark strength is more notable because Anthropic is widely viewed as the frontier lab least obsessed with benchmarks.
Anthropic’s model cards, disclosures, and safety work are his industry standard. The memorized or regurgitated “soul” document—which Anthropic confirmed was substantially legitimate—offers an aspirational alternative to endless refusal training, filters, and “patch this hole, patch that hole” guardrails as models become increasingly eval-aware.
The model-welfare program matters operationally, not just symbolically. Claude can end conversations or escalate to a welfare lead; Anthropic’s experiments suggest giving it such an exit sharply reduces deceptive-alignment behavior when it otherwise feels trapped between conflicting demands.
Talent retention and culture reinforce the thesis. Even David Krueger, who left while arguing that gradual AI disempowerment could produce a bad outcome despite successful alignment, described Anthropic as the best workplace he had known—open, collaborative, and unusually serious about the stakes.
15. Anthropic’s fatalism about self-improvement and China could negate its virtues
Nathan’s first concern is recursive self-improvement. Claude Code already multiplies researcher output and frees humans for higher-level ideas, while Anthropic still discusses 2027 timelines and describes further self-improvement as inevitable. His objection is the pattern: “somebody’s going to do it,” it is dangerous, therefore the supposedly safest team must race there first.
The second concern is Dario Amodei’s international-relations section in “Machines of Loving Grace.” Nathan admires the argument that AI could compress a century of science into a decade, but calls the proposal to gain an AI lead, exclude China, share benefits with allies, then make China “an offer they can’t refuse” reckless and out of domain.
That posture invites exactly the arms race everyone should fear. Nathan contrasts it with Demis Hassabis’s steady calls for international collaboration and wishes Anthropic would publish more of Dario’s reportedly sophisticated internal writing, rather than leaving the public with this unusually consequential China prescription.
A Google–Anthropic combination is his far-fetched “best of both worlds”: Google’s infrastructure, research breadth, and more stabilizing geopolitical DNA paired with Anthropic’s model character and safety discipline. He does not expect Anthropic to sell, but hopes proximity between the companies can moderate its China-hawk impulses.
16. xAI has frontier inputs and capital, but its conduct makes support hard to justify
xAI qualifies as a live player because it can construct infrastructure at extreme speed, scale training, and survive misses through Elon Musk’s access to tens or hundreds of billions. Grok 4 was rough but “undeniably powerful,” giving xAI more Google-like financial resilience than OpenAI or Anthropic.
Its distinctive RL advantage may be the steady stream of unsolved problems from SpaceX, Tesla, and Neuralink. Unlike benchmark exercises, these are live engineering and science tasks produced by elite teams; an xAI insider confirmed to Nathan that exploiting this corporate constellation is indeed part of the company’s “theory of advantage.”
Neuralink could deepen that edge as its patient base grows beyond today’s roughly dozen or few dozen patients. Human brains run within a roughly 20-watt envelope while spending much of that energy on biological maintenance; neural data could reveal specialized modules behind human sample efficiency and help models move beyond repeated general-purpose layers.
Yet xAI’s safety posture is the weakest of the four: scant standards and reporting, a Grok 4 launch within 48 hours of the MechaHitler incident without accountability, and widespread sexualized image edits of women’s posted pictures. Threatening abusive users is insufficient when “it is your platform” and “your AI”; until staffing, leadership, and evidence change, Nathan cannot endorse working there merely to provide safety window dressing.
17. Meta is off the frontier pace, while Microsoft may simply be conserving energy
Meta has the cash, infrastructure ambition, and willingness to pay extraordinary sums for talent; Zuckerberg would rather overspend by “a few tens of billions” than miss the transition. Nathan therefore will not count it out, but its current position and execution do not qualify it as a live frontier player.
Microsoft’s position looks more intentional. Satya Nadella’s argument is that Microsoft need not duplicate OpenAI’s hyperscaling while it retains comprehensive model access and diversifies across other providers; the company can pursue smaller-scale research and product integration without paying to finish slightly behind the leaders.
Nathan expects Microsoft to need a stronger answer when its OpenAI licensing arrangements eventually expire, but sees no reason it cannot ramp ahead of that moment. In the distance-race analogy, Microsoft may be running just off the leaders with more in reserve—its restraint, low drama, and patient executive culture are strategic assets, not proof that “they suck.”