Google Invests $40B Into Anthropic, GPT 5.5 Drops, and Google Cloud Dominates | EP #252
Summary
The frontier-model race is narrowing to OpenAI, Anthropic and Google in the West, while Chinese open-weight labs keep resetting the cost floor. The panel counted 15 major releases in eight weeks and saw the contest shifting from model weights toward inference-time reasoning and compute. As Alexander Wissner-Gross put it, this is an “arms race or horse race or rat race,” with American open models described as roughly six months ahead of Chinese open-weight models.
Kimi K2.6 makes orchestration—not loyalty to one model—the practical winning architecture. The trillion-parameter open-weight model activates 32 billion parameters at once, runs 300 parallel agents, handles text, images and video, and reportedly trained for $4.66 million; Dave Blundin estimated one-eighth the API cost of leading closed models through Fireworks AI and one-thirtieth when self-hosted. His preferred stack keeps a smarter Opus 4.7 orchestrator above cheaper workers, because “your orchestrator will tell you, hey, this is garbage.”
GPT-5.5’s consequential leap is autonomous execution, with research mathematics providing the longer-term clock. Released seven weeks after GPT-5.4, it reportedly uses 40% fewer tokens, cuts hallucinations 60%, improves million-token recall and posts its largest gain on Terminal-Bench 2.0—evidence that OpenAI is fortifying Codex. Wissner-Gross saw FrontierMath Tier 4 improving about one percentage point per month and concluded: “Math is cooked,” even though API prices doubled to $5 per million input tokens and $30 per million output tokens.
Google is assembling a vertically integrated compute position spanning TPUs, NVIDIA systems, cloud capacity and strategic equity. Its eighth-generation TPU 8T and 8i promise three-times-faster training and 80% better performance per dollar, while the committed 960,000-GPU Vera Rubin A5X system is described as larger than Colossus 2 and Stargate Abilene. With Google processing 16 billion tokens per minute and AI writing 75% of its code, Diamandis’s call remained: “Google’s the winner in the long run here.”
Anthropic’s discounted financing reveals that scarce compute can command more bargaining power than headline valuation. Google offered $10 billion now at a $350 billion valuation plus $30 billion contingent on performance and 5 GW of TPU capacity; Amazon’s total commitment reaches $33 billion in exchange for at least $100 billion of AWS spending and another 5 GW. Against a cited $1 trillion secondary valuation, both hyperscalers are effectively buying equity at roughly one-third the price—but the panel argued the binding constraints remain powered land, energy and ultimately TSMC-class fabrication.
The application-layer metric is becoming economic value per token, not raw model prestige or token consumption. Anthropic’s coding and marketplace experiments were framed as attempts to make each token produce more valuable work, while reusable “skills” could become defensible vertical IP. Blundin wants companies moving toward token budgets equal to 50%—perhaps 100%—of payroll; Alex Hormozi’s correction was that reduced iteration time and output matter more than a vanity spend number.
Ambient AI creates an investable collision between assistance, surveillance and verified identity. Chronicle’s constant screenshots could make agents “telepathy-like,” but Wissner-Gross called remote capture an “architectural atrocity” that belongs inside secure local hardware and the operating system. Deepfake losses were presented as rising from $130 million across 2019–23 to a projected $40 billion in 2027, making World ID, cryptographic provenance and authenticated cameras more valuable precisely because synthetic media is improving.
Healthcare supplied the episode’s strongest evidence that AI’s economic impact will extend beyond software. ChatGPT for clinicians reportedly scored 59 on HealthBench versus 43.7 for human clinicians, while AI-assisted donor-heart selection could add 500 transplants and personalized mRNA and CAR-T trials showed durable cancer responses. The disagreement was whether clinicians receive a cognitive exoskeleton or are ultimately replaced; Wissner-Gross said, “Of course, this is about replacing doctors” through an eventual end-to-end system.
Deep dive
1. Model cadence has turned the frontier into a compute race
Peter Diamandis counted 15 major releases in eight weeks—roughly two per week—and noted that the visible field is overwhelmingly American and Chinese. His suspicion was competitive staging: some models may already be “cooked,” waiting to be released directly on top of a rival’s announcement.
Wissner-Gross narrowed the Western frontier to OpenAI, Anthropic and Google, calling it an “arms race or horse race or rat race.” Chinese open-weight models may remain several months behind, but their price-performance keeps pressure on American closed systems; Europe, India and Japan appear to be spectators largely because the US and China hold the compute.
The deeper possibility, attributed to OpenAI’s Noam Brown, is that model weights matter less as inference-time reasoning expands. Wissner-Gross described the result as a “space-time transformer” unfolding through reasoning tokens: if weights commoditize, the decisive variable becomes how much compute a lab can deploy at inference.
Ismail’s summary was broader than benchmark leadership: “The cost of cognition, coordination, and execution [is] all collapsing at the same time.” Wissner-Gross added that current capabilities have outrun even optimistic predictions from three, six, nine or 12 months earlier. Diamandis argued that model self-improvement makes cold-start catch-up harder.
2. The practical consumer unlock is autonomous orchestration
Wissner-Gross initially challenged the consumer framing: OpenAI had bet heavily on consumers consuming vast quantities of reasoning tokens, then pivoted prominently toward enterprise demand. In his view, the near-term economic question is less what an average consumer wants than what enterprises will pay agents to execute.
Blundin’s answer for ordinary users was simpler: choose one of the latest models and “ask it to install itself.” A coordinator can now manage dozens or hundreds of subordinate models, configure systems, download software and explain its work without requiring the user to understand a Linux command line.
That changes software creation more than another benchmark point does. Someone who has never programmed can now think of a product and produce it within an hour; meanwhile, context-window mismanagement elsewhere in the stack still wastes factors of five or 10, leaving room for orchestration and abstraction-layer companies.
3. Kimi K2.6 resets open-weight economics without erasing risk
Diamandis presented Kimi K2.6 as a trillion-parameter open-weight model activating 32 billion parameters at a time, coordinating 300 agents and natively processing text, images and video. Moonshot AI reportedly trained it for $4.66 million and has approximately $4.7 billion of backing from Alibaba, Tencent and IDG.
Blundin’s cost comparison was the tradeable point: running Kimi K2.6 through Fireworks AI costs roughly one-eighth as much as leading Claude or OpenAI APIs, while local deployment could approach one-thirtieth. At scale, where an application may run 10 or 100 agents concurrently, “one-eighth is a pretty damn big price cut.”
Wissner-Gross retained the frontier distinction: Kimi K2.6 and DeepSeek V4 are compelling for privacy, fine-tuning and self-hosted enterprise workloads, but he does not regard them as frontier-equivalent. Blundin countered that Kimi’s benchmarks approached Opus 4.6 only three months after its release, suggesting a three-month rather than six-month lag.
Security remains the discount’s hidden cost. No human can inspect millions of AI-generated code lines, so the panel’s honest answer to prompt or code injection was that “you have to use AI to protect against AI”; a stronger closed-model orchestrator can review cheap workers, but it cannot provide an absolute guarantee.
4. Sparse experts explain how trillion-parameter models stay economical
Diamandis used Kimi to unpack mixture-of-experts routing: instead of activating a trillion parameters for every token, an orchestrator selects the subsets relevant to coding, mathematics or another domain. That lowers both latency and cost without shrinking the model’s total stored capability.
Wissner-Gross placed mixture of experts within the wider principle of sparsity, which he said is endemic across frontier models and analogous to the human brain’s selective firing. Sparsity also regularizes models by preventing individual weights from overfitting the training data.
David Friedberg added that routing can happen through roughly 140 layers, selecting among as many as 128 experts at each layer rather than merely choosing one top-level persona. Wissner-Gross’s “holy grail” is eventually a million-parameter-or-smaller “diamond or black hole of a model,” with increased sparsification as one possible route.
5. GPT-5.5 strengthens Codex and starts a research-math countdown
GPT-5.5 arrived seven weeks after GPT-5.4 with native text, audio, image and video handling. The reported gains included a 37-point improvement in long-context reasoning, more reliable use of its million-token window, 40% fewer tokens at equal latency and 60% fewer hallucinations.
Wissner-Gross focused on Terminal-Bench 2.0, where the largest jump measures autonomous command-line operation. Having used the model, he did not see mere benchmark overfitting, but the release “very much feels like” an effort to make Codex a stronger Claude Code competitor.
FrontierMath Tier 4 supplied his more radical conclusion. GPT-5.5 Pro gained roughly two percentage points in two months and approached solving half the professional research-grade problems; at even a flat one-point monthly pace, the remainder would fall in four or five years. “Math is cooked. Bunch of other things are cooked.”
Users pay for that gain: API prices doubled from $2.50 to $5 per million input tokens and from $15 to $30 per million output tokens. Blundin nevertheless found the execution gain tangible—complex integrations “just flat out work”—while Demis Hassabis’s estimate of remaining conceptual breakthroughs has moved from five more than 10 years ago toward a coin flip that scaling alone suffices.
6. Google’s full-stack compute position is widening
At Google Cloud Next 2026, Google unveiled eighth-generation TPU 8T for training and TPU 8i for inference, claiming three-times-faster training and 80% better performance per dollar. The chips are explicitly designed for running millions of agents in real time.
David Friedberg found it striking that Google’s systolic-array TPUs land near NVIDIA on price-performance despite their different architecture. Matching NVIDIA is enough when Google owns the stack from chip design and data centers through cloud and models; Diamandis said, “I still believe Google’s the winner in the long run here.”
Sundar Pichai said Google was processing more than 16 billion tokens per minute and AI now writes 75% of Google’s code. A cited estimate put Google at roughly one-quarter of planetary AI compute, while Wissner-Gross noted that Google AI is already helping design the next TPUs—“recursive self-improvement goes all the way down to the silicon.”
Google also committed to 960,000 NVIDIA Vera Rubin GPUs for its A5X bare-metal instance, promising 10-times-higher token throughput and 10-times-lower inference cost. The system was described as twice the scale of Colossus 2 and 2.4 times Stargate Abilene, complementing rather than replacing Google’s proprietary silicon.
7. Anthropic is exchanging discounted equity for power and chips
Google’s proposed $40 billion package consists of $10 billion now at a $350 billion valuation, another $30 billion conditional on performance, and 5 GW of TPU compute over five years. Diamandis contrasted that valuation with approximately $1 trillion on secondary markets: strategic compute providers are buying at roughly one-third the outside price.
Amazon’s parallel deal brings its Anthropic investment to $33 billion—$25 billion beyond the previous $8 billion. Anthropic, in turn, commits to spend at least $100 billion on AWS over a decade, run Claude on Trainium and consume another 5 GW of Amazon capacity.
Blundin relayed, with an explicit anonymous-source caveat, that Anthropic internally might reach $40 billion to $50 billion—and perhaps $70 billion—of revenue by year-end. The reason it might miss is not demand but compute; the panel linked Mythos’s limited release, and possibly OpenAI’s Sora pullback, to the same constraint.
Diamandis’s blunt formulation was: “Dario needs compute.” Yet all the cross-investment ultimately meets foundry capacity at Samsung, Intel and TSMC; Blundin said that $16 billion, potentially $45 billion, of Samsung capacity had been locked up. Wissner-Gross judged powered land and permitting more binding today, with semiconductor supply likelier to become the medium-term stranglehold.
8. Economic value per token will decide the application layer
Blundin’s infrastructure call was to back operators that possess both chips and power. Legacy industrial loads such as aluminum smelting can be converted into far more valuable data-center capacity, while kernel-level software that unlocks AMD, older GPUs or more efficient NVIDIA inference can monetize scarcity without building a model.
At the enterprise layer, Anthropic’s “skills” let agents pull reusable procedures into context instead of rediscovering them. Blundin expects companies to refactor around hundreds or thousands of such skills, turning a complete vertical workflow library into “defensible intellectual property.”
Anthropic’s Project Deal, which had Claude buy, sell and negotiate in an internal marketplace, briefly invited comparisons with eBay. David Friedberg saw the stock reaction as knee-jerk; Alex Salkever’s sharper concern was that AI makes coordination easy across listings, support and disputes, allowing entire workflow categories—not necessarily incumbents—to be automated.
Wissner-Gross resolved Anthropic’s strategy with one rule: “They’re trying to maximize the economic value per token.” Code generation won because useful code is worth more per token than consumer video or cat images; marketplace and business-running experiments are searches for the next similarly valuable output.
9. The Musk–OpenAI trial matters even without a decisive verdict
The episode recorded as jury selection began in Oakland federal court for Elon Musk’s case against Sam Altman and OpenAI. The panel expected discovery to expose private texts and emails and compared the conflict’s cultural significance with the Apple–Microsoft battles that later became docudramas.
Wissner-Gross described a two-stage structure: first determine whether Musk’s claims hold, then consider remedies. He also noted reports that political views of Musk could seep into jury selection and that the judge might treat the verdict as advisory while determining any final award from the bench.
Blundin’s strategic point was that Musk need not fully win: “He just has to slow down OpenAI.” In a frontier race moving month by month, a three-month delay could itself be a loss, regardless of whether invested capital is eventually impaired or OpenAI’s corporate conversion is reversed.
10. Ambient AI is useful because it watches—and dangerous for the same reason
OpenAI’s Chronicle periodically captures a user’s screen, sends images to servers for OCR and visual analysis, then builds structured local memories. Altman called the resulting experience “telepathy-like”; Diamandis saw the same mechanism as both an expert over the shoulder and the camel’s nose of worker replacement.
Wissner-Gross rejected privacy loss as inevitable. He called server-side screenshot capture an “architectural atrocity” that should be native to the operating system, compositor and secure enclave, with hardware guarantees that rendered pixels and derived memory never leave the device.
Wissner-Gross argued that screen awareness and voice are the two missing interfaces for mainstream adoption. Instead of asking a user to navigate layered settings, an agent can say, “I see what you’re doing wrong. In fact, let me just do it for you”—the same feedback loop developers already create manually with screenshots.
Workplace deployment will be contentious. Blundin cited a statistic that 44% of Gen Z workers are sabotaging automation by feeding systems bad training data; he called that a “losing battle,” while Diamandis openly accepted the opposite bargain—giving an agent nearly all personal data to maximize usefulness.
11. Synthetic media makes verified humanity more valuable
Diamandis traced deepfake losses from $130 million across 2019–23 to $400 million in 2024, $1 billion in 2025 and a projected $40 billion by 2027. The emblematic case was the Arup incident in Hong Kong, where an employee authorized $25 million of transfers after a video call populated entirely by synthetic colleagues.
World ID’s Zoom integration combines Orb enrollment with real-time face authentication to place a verified-human badge on participants. Alex Hormozi saw human identification, rather than World’s crypto economics, becoming the stronger product: “The more AI scales, the more valuable verified human identity becomes.”
David Friedberg supplied the personal specimen: his controller believed an impersonator’s urgent instructions and sent $300,000 toward China; roughly $75,000 crossed the border and was never recovered. He expects identity, logging and transaction controls eventually to make digital fraud more tractable than chemical, biological or radiological threats.
A Grok-generated French woman holding a convincingly reflective ID showed why photographed documents will not suffice. One panelist suggested centralized verification; Diamandis advocated hardware cryptography and chain-of-custody for cameras, and noted a bipartisan House bill addressing deepfake fingerprinting.
12. Token-maxxing is diagnostic, but productivity remains the real metric
A 404 Media report described startup CEOs bragging that AI compute cost more than equivalent human labor. Blundin rejected the backlash framing: the immediate risk is failing to experiment, while inefficient token use can be optimized after an organization learns what agents can do.
His year-end target was tokens equal to 50% of payroll, though he suggested one-to-one may be better if tokens deliver a 10-times force multiplier. The analogy was a salesperson’s miles flown: terrible as a productivity measure, but an answer of zero can reveal insufficient activity.
Alex Hormozi’s correction was to measure tokens against compressed iteration cycles and efficiency, not raw spend. That matched Wissner-Gross’s economic-value-per-token framing and separated genuine operating leverage from a vanity metric.
Wissner-Gross has already seen human-labor-to-AI-compute allocations around 1:2, with frontier labs more asymmetric. His endpoint was “one to infinity effectively”—all tokens and no service labor—prompting David Friedberg’s escape clause: “until the humans merge with the tokens.”
13. The UAE is treating agentic government as national infrastructure
Sheikh Mohammed’s announced target is for agentic AI to run 50% of UAE government sectors, services and operations within two years. His formulation went beyond copilots: AI “analyzes, decides, executes and improves in real time” as an executive partner.
Ismail credited the speed to concentrated authority and alignment. As a test case, officials were challenged to issue his golden visa within five hours rather than Singapore’s cited five days; after initial alarm, they completed it.
Salim offered the example of a western-state wind-turbine approval process in which mapping power lines, water mains and flight paths reportedly reduced approval from six months to roughly 30 seconds. Blundin doubted Congress could copy the UAE’s model, while Diamandis was unexpectedly more optimistic.
14. Clinical AI is moving from assistance toward full-stack medicine
ChatGPT for clinicians was presented as a free copilot for US physicians, nurses and physician assistants. OpenAI reported a HealthBench score of 59 versus 43.7 for human clinicians, validation across 700,000 model responses and 99.6% accuracy under physician evaluation.
Diamandis predicted it will become malpractice to diagnose without AI in the loop. His Fountain Life example generates 200 gigabytes of genomic, imaging, microbiome, metabolic and blood-biomarker data per patient—far beyond what one clinician can integrate unaided.
Wissner-Gross saw the free product as a reference design and distribution channel into biomedical enterprises, where the larger revenue pool lies. OpenAI’s internal benchmarking also points toward comparable products for law, management consulting and finance, even as clinical AI already faces incumbents such as Epic-linked tools and OpenEvidence.
The disagreement was explicit. Jent described the ideal as a cognitive exoskeleton for clinicians, while Wissner-Gross replied, “Of course, this is about replacing doctors,” nurses, HMOs and drug-design workflows end to end. Wissner-Gross also forecast fierce regulatory resistance, while noting that clinicians may embrace AI as relief from hated EMRs—“until it takes their job.”
15. AI can stretch organ supply before biology makes donors obsolete
About 4,000 patients currently need hearts and 103,000 need some form of transplant, yet only one-third of available hearts are selected. A surgeon may have 15 minutes at 2 a.m. to judge viability from a few familiar signals.
NYU and Stanford’s TopHeart examines 20 variables and aims to add roughly 500 hearts to the transplant pool by giving that surgeon a second opinion. Wissner-Gross paired smarter selection with national matching, vitrification and cryopreservation as ways to improve today’s constrained distribution system.
The panel still viewed donor optimization as transitional. Xenotransplantation, bioprinting and efforts to grow a patient’s heart, liver, lung or kidney from reprogrammed skin cells could create abundance by decade-end; Wissner-Gross’s preferred endpoint is one where another person never has to die for a transplant to occur.
16. Personalized immunotherapy is becoming operational medicine
In the pancreatic-cancer trial presented, historical five-year survival was 13%; eight of 16 patients mounted a strong vaccine response, and 87.5% of those responders were alive after six years. The approach sequences a removed tumor, selects 20 mutations and builds a personalized mRNA vaccine that directs killer T cells toward the cancer.
More than 120 mRNA cancer-vaccine trials were cited across lung, breast, prostate, melanoma, pancreatic and brain cancers. Wissner-Gross described the platform as a universal solution in spirit; he joked that the long-promised medical nanobots arrived as lipid nanoparticles: “They’re just fat.”
A separate melanoma CAR-T result was presented as all 20 patients becoming minimal-residual-disease negative within two months, with 15.3 months’ median follow-up. Diamandis used “cure”; Wissner-Gross celebrated the result but called blood extraction and ex-vivo cell engineering “horse and buggy era,” asking when the same reprogramming becomes autonomous and in vivo.
Drug repurposing supplies a cheaper parallel path. Candesartan, an approved blood-pressure medicine, was presented as active against MRSA, while David Fajgenbaum’s use of rapamycin against his own Castleman disease inspired Every Cure’s AI search across roughly 18,000 diseases and 4,000 approved drugs. The old liability of “off-target” effects becomes a searchable therapeutic asset.
17. Embodied AI is crossing from novelty into transport economics
The ACE table-tennis robot used nine cameras and three vision systems and won three of five games in the demonstration. Wissner-Gross was surprised such a low-dimensional task took this long; Blundin’s inference was that cheap transformer vision and feedback now let one or two people solve many adjacent home-robot tasks in weeks.
Tesla’s Cybercab entered production with no steering wheel or pedals, an advertised $30,000 price and operating cost of $0.20 per mile. Its two seats fit the cited 1.2-person average Uber load, while Tesla’s stated ambition is two million units annually; Diamandis imagined owners deploying fleets that earn while they sleep.
The manufacturing thesis is radical simplification: Diamandis contrasted roughly 2,000 drivetrain components in an internal-combustion car with 17 moving drivetrain parts in a Tesla. Shorter range also matters less for an autonomous fleet because a car can recharge itself while another answers demand.
Joby’s demonstrated JFK-to-Manhattan air-taxi trip compressed a 16-mile, hour-plus drive into seven minutes, with four passengers, one pilot and noise claimed at one-hundredth of a helicopter’s. The panel’s “why did this take so long?” answer was permitting; Diamandis reframed every unresolved robotics bottleneck as a near-term job, while the closing response was that AI models would attack those bottlenecks too.
18. Work, education and ownership must all be refactored together
Ismail said Accenture- or Capgemini-style consulting is in “very big trouble” if it remains a pyramid of junior analysts producing decks. Survivors become intelligence platforms combining domain expertise, agentic workflows, benchmarks, governance, implementation and change management—not firms renting headcount by the hour.
Wissner-Gross rejected the idea that humans must originate every venture. AI can propose and operate hundreds or thousands of microbusinesses, while the owner supplies taste: “Do you like it? Yes. No.” Diamandis agreed that ideas were rarely the bottleneck; execution was, and AI is rapidly removing it.
Blundin argued that eliminating entry-level jobs is not like eliminating babies because apprenticeship-style career ladders were already mismatched to singularity-speed change. AI becomes the teacher, training becomes nimble, and engineering degrees may ultimately be awarded for “what did you build?” rather than courses completed.
Diamandis’s career advice was to “build in public”: a GitHub portfolio becomes the résumé, demonstrated output replaces pedigree, and capable builders may choose entrepreneurship over employment. He still expects a human-interface layer because people value other people, but acknowledged that AI can eventually perform any job.
Wissner-Gross identified the unresolved political economy: if productivity explodes while income remains concentrated, consumer demand collapses because “capitalism needs customers.” AI dividends, wider equity ownership, sovereign funds and drastically lower costs are institutional choices, not automatic consequences of technology.
On existential risk, Blundin cited estimates of 10–20% from Musk and Hinton, 25% from Dario Amodei and roughly 10% from Altman in an interview. His explanation was not acceptance but race logic: each lab trusts itself and believes a pause would leave China moving, while government’s continued inaction is “utterly insane.”
Wissner-Gross questioned whether those public probabilities match revealed behavior. He separately argued that AGI, under his definition, arrived no later than summer 2020—roughly nine years ahead of Ray Kurzweil’s 2029 date—and that the singularity is not a point in 2045: “It’s now and it’s an interval and we’re right in the middle of it.”