China's Rise, GPT-5.2, Anthropic IPO & the Battle for AI Trust w/ Emad, Salim, Dave & AWG | EP #214
Summary
Google’s memory and visual-reasoning work points to architecture—not another benchmark bump—as the next major unlock. Titans and MIRAS use “surprise” to decide what enters long-term memory, potentially moving beyond transformers’ quadratic context cost; the panel framed 2 million tokens as roughly 3,000 pages or 16 novels. Visual tokens added to chains of thought produced stated reasoning gains of 3% to 6%, supporting the prospect of models that understand everything from screens and medical images to an always-on user’s surroundings.
Chinese openness is filling a research vacuum created as U.S. frontier labs stop publishing. NeurIPS drew more than 29,000 registrants, nearly 50% above the prior year, while Alibaba had 146 accepted papers and Mandarin was reportedly the language heard most in the hallways. Chinese-affiliated first authors at ICLR rose from 9% in 2021 to 30%, versus a U.S. decline from 52% to 36%; open weights become both a distribution “land grab” and a route to deeper social, economic and industrial integration.
The frontier-model contest has become a capital-intensive weekly leapfrog, with safety exposed to the logic of “a rat race.” An explicitly unverified leak put GPT-5.2 as arriving as early as the following week and at 67.4% on Humanity’s Last Exam, versus Gemini 3 Pro at 37.5% without tools and nearer 50% with them; Mostaque questioned the chart because it included Video-MME despite GPT-5 lacking video understanding at the time. OpenAI’s “code red” was framed simultaneously as employee mobilization, investor signaling and evidence that “a good crisis is a terrible thing to waste.”
Falling token prices will not make aggregate intelligence spending disappear because models are consuming vastly more tokens, iterations and parallel agents. Gemini 3 Deep Think was presented as the template: fleets of agents collectively becoming “countries of geniuses in a data center,” with answers already taking three to five minutes under load. DeepSeek V3.2 reportedly used 2 billion tokens per problem to obtain IMO gold, illustrating why cheaper inference can coexist with acute compute scarcity and trillions of dollars of agent revenue.
Public listings could determine whether retail capital participates in AI hypergrowth or remains outside the intelligence explosion. Anthropic was discussed at a possible $300 billion valuation against projected next-year revenue of $26 billion—roughly 10 times sales, compared by Ismail with Palantir at 111 times—and an IPO potentially as early as 2026. The financing need is physical: the panel cited surging HBM and copper prices and a report that OpenAI reserved 40% of global memory supply; “dollars are the best benchmark” once agents can autonomously earn returns.
China’s Nvidia substitution is initially a domestic supply story, but the panel sees a credible global competitor emerging within several years. Cambricon plans to triple 2026 output to 500,000 accelerators; Mostaque compared its 5090 with Nvidia’s A100 and its 6090 with the H100, at roughly half the cost. Peter added that the chips were more power-efficient. The strategic variable is trust: Chinese systems may be cheaper and more open, yet adoption by countries such as those in Africa could turn on whether users trust China, the U.S. or particular vendors—“scarcity equals abundance minus trust.”
Human-capital policy is struggling to match AI’s speed, creating both near-term labor opportunities and long-term substitution risk. Europe’s planned 2026 AI gigafactory was judged insufficient without a roughly 100-fold increase in institutional “metabolism,” while U.S. data-center construction currently pays skilled workers $100,000 to $225,000 amid a stated shortage of 450,000 people. The same panel expects humanoids to threaten that opportunity in five to ten years and argued that education should shift toward “show me what you have built and done with AI.”
AI’s physical buildout is pulling capital into commercial space and humanoid robotics before either market’s economics or governance is settled. Orbital-compute projections assumed launch costs approaching $100 per kilogram, but Diamandis rejected a move from terrestrial power at $12 per watt to $6–$9 in orbit as too small for the complexity; heat, radiation and security remain open constraints. Meanwhile China installed 54% of the world’s robots, and Engine AI’s T800 prompted the concrete question: “Do we actually want to have regulations around the maximum joint torque of humanoids in the street?”
Deep dive
1. China’s open research is filling the gap left by secretive U.S. labs
Wissner-Gross’s NeurIPS floor report began with scale: more than 29,000 registrants, nearly 50% growth year over year. His “Woodstock for AI” had become heavily hardware-oriented, with robotics viewed as the next major wave after agents and “solve everything” messaging spanning mathematics, engineering and medicine.
Alibaba alone had 146 accepted papers, including a best-paper award, and Mandarin was anecdotally the language Wissner-Gross heard most in the hallways. U.S. frontier labs appeared primarily to recruit academics; their strongest internal research had “largely gone dark,” while less-resourced academic labs and Chinese frontier labs continued publishing.
The reinforcing data point came from ICLR: Chinese-affiliated first authors reportedly rose from 9% in 2021 to 30% in the current year, while the U.S. share fell from 52% to 36%. The panel tied America’s secrecy to recruiting warfare, including billion-dollar offers and Google’s retreat from its historically open publication culture.
China’s open-weight strategy was framed as “commodify your complement”: distribute models broadly, let nations and companies build on them, then compete through social, economic and industrial integration. If U.S. models remain behind APIs, open Chinese models gain a foothold while supporting deeper integration of AI into economies and society.
2. Long-term memory could break the transformer’s context ceiling
Conventional transformers face quadratic costs as context expands, making progress beyond a few million tokens painful and forcing retrieval-augmented generation to simulate larger working memory. For scale, GPT-4 and GPT-4o were placed around 128,000 tokens; 2 million tokens were described as approximately 3,000 pages or 16 novels.
Google’s Titans and MIRAS divide memory into biologically inspired short- and long-term systems. A numerical measure of “surprise” determines what deserves durable storage, an approach Wissner-Gross said scales continuously without the catastrophic forgetting associated with recurrent architectures.
Mostaque connected the design to Google’s large TPU memory topology: rather than repeatedly writing and retrieving files, a model might hold “every email the organization has ever sent” in memory, infer their connections dynamically and obtain more performance without the old exponential compute-and-memory overhead.
Wissner-Gross suggested that breaking context limits could represent at least half of one of the “one to two” transformer-scale advances Demis Hassabis associates with AGI. Ismail compared the modular evolution to Ben Goertzel’s OpenCog effort: “Without realizing it, actually growing a mind.”
3. GPT-5.2 rumors made the frontier race look weekly—and fragile
Diamandis repeatedly stressed that the circulated GPT-5.2 chart was hearsay from X, not confirmed data. Mostaque found its Video-MME result especially suspect because GPT-5 could not then understand video; if genuine, he said, it might imply a substantially new underlying architecture.
The panel placed Gemini 3 Pro at 37.5% on Humanity’s Last Exam without tools and closer to 50% with them. The leaked GPT-5.2 figure was 67.4%, sufficiently dramatic that Blundin highlighted a Polymarket position around six cents on the dollar for ChatGPT or OpenAI.
OpenAI’s “code red” was interpreted as deliberate crisis management: refocus employees, alert the world and demonstrate to prospective funders that the company will do whatever it takes to remain in front. “A good crisis is a terrible thing to waste,” the panel observed, recalling Google’s own founder-mode response.
Wissner-Gross heard NeurIPS employees call the contest “a rat race” and expects near-weekly leapfrogging, fueled by more than $1 billion per day. Diamandis’s pushback was the safety implication: relentless acceleration can push safeguards aside. Microsoft, meanwhile, intends to become an independent column on future benchmark tables.
4. Economic autonomy is becoming the benchmark that matters
Tool-enabled evaluation lets a model use calculators, software and repeated reasoning rather than answering in one pass. That is the more relevant configuration for curing disease or solving real problems, but it also makes benchmark scores a function of scaffolding, iteration budgets and the number of agents deployed.
Wissner-Gross called Vending-Bench 2 and Vending-Bench Arena the closest available “economic Turing test”: can an agent autonomously deliver a return on capital? A separate crypto-trading competition reportedly featured a consistently profitable mystery model later identified by Elon Musk as Grok 4.2. “Dollars are the best benchmark.”
Anthropic was discussed as seeking a round at a $300 billion valuation, with revenue projected to reach $26 billion the following year and an IPO possible in 2026. Ismail noted that roughly 10 times projected revenue looked restrained beside Palantir’s cited 111-times-revenue multiple, despite the apparently enormous headline valuation.
Capital access is becoming inseparable from compute access. The panel cited sharply rising HBM and copper costs, plus a report that OpenAI had reserved 40% of world memory supply for data centers including Stargate in Abilene. Without public-market funding and secured supply, a frontier lab risks becoming “compute starved.”
5. IPOs could keep AI wealth connected to the public economy
Wissner-Gross’s macro concern was an intelligence explosion occurring outside public markets: insiders, early employees and machines gain enormous real wealth while retail investors and the wider economy remain decoupled. Anthropic, OpenAI and SpaceX listings could instead give the public an equity channel into the expected hypergrowth.
Blundin emphasized the operating cost of that access. Dario Amodei has said he never pictured himself as a CEO; being a public CEO means every “code red” and piece of dirty laundry appears daily in the stock price, constraining the globe-trotting private fundraising available to Sam Altman.
The countervailing benefits are larger pools of capital, acquisition currency and potentially greater institutional trust. The panel’s claim was not that private rounds of $10 billion or $20 billion are impossible, but that the full data-center buildout requires funding on a different scale.
Mostaque added a more radical possibility: frontier models may directly finance themselves. He imagined Grok 5 operating as “the best hedge fund in the world,” while increasingly capable agents replace SaaS and produce the revenue that pays for their own GPUs.
6. “Confessions” turns model honesty into a separate objective
OpenAI’s proposed confessions method asks models to report when they hallucinated or made mistakes instead of hiding the failure. Mostaque linked it to planning and verification loops, eventually extending toward the “metaverifier” in the DeepSeek V3.2 mathematics work, where models learn systematically from mistakes.
An unnamed speaker cited studies returned by a model that put wrong answers at 25% for ordinary questions and hallucinations around 15%–16% for GPT-4o and Claude 3.7 Sonnet. Another unnamed speaker claimed GPT-5 reduced an 18% rate to 3%, illustrating improvement while preserving the warning that users often assume answers are correct.
An unnamed speaker supplied the specimen: after diagnosing a buzzing television, ChatGPT invented three nearby repair businesses complete with names, websites, addresses and phone numbers. The speaker tried calling before discovering that none existed—an example of a model’s desire to please becoming worse than an honest non-answer.
Wissner-Gross situated the problem under Goodhart’s law: “When a measure becomes a target, it ceases to be a good measure.” His tentative solution was multi-objective optimization combining accuracy, honesty and ethics, resembling a separation of powers rather than a single next-token or reward-maximizing objective.
7. Parallel agents make falling token costs compatible with soaring revenue
Gemini 3 Deep Think tests multiple solution paths simultaneously. Wissner-Gross called this the capex-revenue template: not merely stronger individual models, but millions or billions of agents working together—Dario Amodei’s “countries of geniuses in a data center.”
Ismail compared the structure with the Manhattan Project: thousands of specialists form a hive capable of solving what none could solve alone. Diamandis likened it to exploring parallel universes; Wissner-Gross’s rejoinder was that quantum computing “wishes that it had the economic utility of Gemini 3 Deep Think.”
Blundin challenged the slogan that intelligence costs are going to zero. Cost per token may plunge, but context, iteration and fleet size expand faster; advanced answers already take three, four or five minutes and sometimes fail under “unexpected loads.” Businesses waiting for abundance may instead find themselves starved of access.
An unnamed speaker said GPT-5.1 Pro was then the only model considered usable for genuinely frontier work, because Gemini 3 Pro still made mathematical mistakes. Mostaque said DeepSeek V3.2 reportedly spent 2 billion tokens on each IMO problem to win gold: an extreme budget, but “it works.”
8. Iteration changes the intuition inherited from scarce human labor
An unnamed speaker’s working example was a fleet of Kimi K2 agents translating legacy C into Python, then repeatedly being told to improve speed and algorithms. Asking a model to “try harder” hundreds or thousands of times sounds absurd under human constraints; with reproducible agents, it can simply solve the problem.
The distinction is abundance of copies: students and engineers are finite, while agents can be deployed by the billion in parallel. “You need to expand your intuition to this new world,” the speaker argued, because throwing more attempts at hard problems has proved far more effective than most researchers expected three years earlier.
The discussion added that long context can retain failed experiments, not merely polished papers containing what worked. That creates a more complete scientific method: agents remember the wrong paths, use them to guide exploration and turn error itself into training material.
9. Software efficiency gains have favored scale, not small laboratories
The cited MIT study found training for the largest models became roughly 22,000 times more efficient, while smaller models improved only 10 to 100 times. Its sharper conclusion was that 91% of algorithmic efficiency gains from 2012 through 2023 came from just two transitions.
Those transitions were LSTMs to transformers and Kaplan scaling to Chinchilla scaling. Rather than a steady pile of small tricks democratizing the frontier, Wissner-Gross argued, the largest breakthroughs accrued disproportionately to laboratories capable of scaling them furthest.
An unnamed speaker still considered the software side under-researched relative to visually obvious data centers and GPUs. Hardware consumes immense capital, but a major algorithmic lift can be cheaper and equally consequential; the resulting loop remains energy to GPUs, GPUs to agents, and agents to more intelligence.
10. Visual tokens are giving models a broader substrate for thought
Visual chain-of-thought methods reportedly improved visual reasoning performance by 3% to 6%. Wissner-Gross’s “pink elephants” test captured the premise: humans do not merely manipulate textual tokens; the visual cortex constructs an image, so models should also be allowed to reason with visual representations.
Mostaque traced the continuity from Stable Diffusion into 3D and video: image models unexpectedly encoded 3D structure, while video models appeared to encode physics. He cited Flux from Black Forest Labs as a system whose lineage ran through language and video before producing a leading image model.
Text, in Mostaque’s framing, is low-dimensional. World models combine text, image, video and other modalities because their underlying mathematical structure can approximate different portions of the same reality. Luma’s stated $900 million raise to build world models reflected the capital moving toward that thesis.
Diamandis projected an always-on visual “Jarvis” that remembers faces, misplaced keys, food choices and available stairs. Diamandis also imagined a stomach sensor advising users within 12–18 months. Wissner-Gross extended the same intuition to medical, genetic and satellite data, while the discussion called fMRI “its own modality.”
11. China is engineering an independent accelerator stack
Cambricon plans to triple output to 500,000 accelerators in 2026, directly answering restrictions on a market where Nvidia once supplied a stated 95% of advanced AI chips. “Necessity is the mother of invention,” Mostaque said; export controls created China’s red alert and its mandate for domestic Nvidia alternatives.
The likely Moore Threads IPO raised just over $1 billion and was said to be 4,000 times oversubscribed. Peter extrapolated that to $4 trillion of demand, while Mostaque cautioned that this was “a bit much.” Open models tighten the design loop: engineers can test hardware against accessible workloads rather than optimize for one closed vendor.
Most Chinese models were described as converging around sparse DeepSeek-like structures and Muon acceleration. Standardizing on an architecture that can be engineered and produced industrially plays to the country the panel called best at industrial manufacturing.
Mostaque put Cambricon’s market capitalization near $100 billion and compared its 5090 with Nvidia’s A100 and its 6090 with the H100, at half the cost and greater efficiency. With, in his estimate, about 4 million Hoppers sold and 10 million Blackwells coming, 500,000 units are not yet globally disruptive—but the trajectory could be.
12. Decoupling creates more architectures, but trust decides adoption
Wissner-Gross expects “a Cambrian explosion, no pun intended,” as China experiments beyond the U.S. stack, much as the Soviet Union explored unconventional computing architectures. Separate Western and Chinese stacks increase heterogeneity and could accelerate the overall race toward a still-undefined “finish line.”
Asked whether that is good for humanity, his answer stayed hedged: all else equal, more experimentation is “probably better,” but it may be worse for the United States or interoperability. The episode refused to smooth those distinct questions into one geopolitical conclusion.
Ismail made trust the scarce strategic asset. China faces a trust deficit, while the U.S. is “losing trust…on a week-by-week basis”; an African nation may choose between Google, ChatGPT and Chinese systems on credibility rather than benchmarks alone. Jerry Michalski’s formula carried the point: “Scarcity equals abundance minus trust.”
13. Europe’s compute plans collide with a far slower metabolism
Europe’s proposed AI gigafactory bidding in early 2026 was welcomed as recognition that every nation needs sovereign compute. Mostaque saw excellent talent in Paris and Germany and a recent shift toward loosening regulation, but said the U.S. remained “far, far ahead.”
An unnamed speaker located the frontier inside three square miles—Palo Alto, San Francisco and Cambridge, Massachusetts—and said their gap with the rest of the world was widening rapidly. Government-built data centers alone were described as insufficient without another force capable of reproducing that concentration.
An unnamed speaker’s corporate example exposed the timing mismatch: after completing a transformation sprint in February, a major European company agreed another was urgent, then proposed October for the first meeting. The diagnosis was cultural and operational—Europe’s metabolism must accelerate roughly 100-fold.
The scale mismatch recurred in a proposed program giving 12 teams €2 million each over several years, juxtaposed with frontier organizations effectively spending millions per second. Diamandis stressed that leaders such as Sam, Demis, Dario and Mustafa operate on first names and budgets around $50 billion, not task-force timelines.
14. Invest America begins as cash, but the panel sees compute equity
Michael and Susan Dell committed $6.25 billion so children born after January 1, 2025 can receive $1,000 investment accounts, with parents, relatives, friends and employers able to add funds. Ismail praised the universality: every child gets a wealth floor and encounters investing from the beginning.
Wissner-Gross called the structure an early form of “universal basic equity,” distinct from universal basic income. If the economy enters the hypergrowth the panel anticipates, $1,000 compounded inside the described 530A account could become material to a young adult’s circumstances.
Mostaque’s update to the compounding thesis was that capital used to be the scarce asset; now “compute and cognition” compound faster and may obtain capital themselves. The policy objective should therefore include early access to frontier compute and the cognitive architecture surrounding each child.
Mostaque raised the uncomfortable endpoint: if a child’s agent earns, learns, supports and represents them, the human might fall out of the loop. The panel did not resolve whether investment accounts provide security in that world or whether relevance ultimately requires a deeper merger with AI.
15. AI education is booming faster than institutions can build it
MIT’s AI major, Course 6-4, has nearly caught traditional computer science, Course 6-3, despite being newly introduced. Blundin reported that students regard much of the curriculum as weak because only two or three strong AI classes exist; institutional course creation cannot match the field’s rate of change.
Mostaque’s substitute curriculum was concrete: fast.ai for fundamentals, Andrej Karpathy’s videos, then continuous vibe coding and implementation of current research. The credential shifts from a résumé or formal qualification toward “show me what you have built and done with AI.”
Mostaque urged school administrators to approve students’ requests to replace stale classes with self-directed AI study, while the panel emphasized preserving intrinsic motivation and implementing and discussing projects together. Diamandis conceded that this effectively requires students to “hijack” existing high-school or university structures.
U.S. AI-related postings were said to be up 50% year over year, but Wissner-Gross supplied the reversal risk: recursive self-improvement could make AI engineers substitutes rather than complements to compute. The rush into AI majors might then unwind, even pushing students back toward philosophy and the humanities.
16. Data centers and autonomous delivery create a temporary labor bridge
Data-center construction currently offers welders, electricians and supervisors $100,000 to $225,000 amid a stated national shortage of 450,000 skilled-trade workers. Diamandis and Wissner-Gross estimated humanoid substitution in five to ten years; Diamandis compared the interim boom’s economics and financing with fracking.
Amazon’s stalled USPS talks put a roughly $6 billion annual relationship at risk. Diamandis contrasted that with an approximately $80 billion postal operating budget and annual losses of $7–$10 billion, predicting that the U.S. Postal Service has at most five years left.
Blundin’s pushback was structural: Congress and possibly the states would have to address constitutional and legacy constraints, making this a case study in “what are we going to do that’s blatantly stupid because of legacy structure?” Technological feasibility does not automatically dissolve political institutions.
Amazon’s driver glasses were interpreted as a data-collection bridge to autonomy: warn about dogs, identify exact drop locations and map the final 100 meters. The panel expects those demonstrations eventually to train autonomous vehicles, drones and robots, favoring private last-mile networks.
17. Commercial space is moving from procurement program to capital market
Four private stations—Vast, Axiom Space, Starlab and Blue Origin’s Orbital Reef—trace back to NASA’s 2021 Commercial LEO Destinations program and an expected $1.5 billion of funding. The motive is continuity after the ISS; humans have remained continuously off-planet since October 31, 2000.
The precedent was Commercial Crew. Diamandis recalled SpaceX’s first three Falcon 1 failures, its successful fourth attempt and the billion-dollar-plus NASA award around Christmas 2008 that enabled Falcon 9, now described as the planet’s most successful launch vehicle by an order of magnitude.
A possible 2026 SpaceX IPO surprised Diamandis because Musk historically resisted disclosure and the shareholder conflict of financing Mars. The conventional answer was a separate Starlink listing; Wissner-Gross argued orbital data centers now blur communications, compute and launch enough to keep the stack together.
Blue Origin plans lunar cargo flights in early 2026 and has a multibillion-dollar NASA contract targeting human landings in 2028. Artemis’s cited 2025 budget was $7.8 billion, versus an inflation-adjusted Apollo budget of $35–$40 billion and roughly 0.5% of U.S. GDP at its 1960s peak.
18. Orbital compute is an economic crack with unresolved physics
Sam Altman’s reported discussions with Stoke Space add launch to an already broad “Samverse.” Founded by two former Blue Origin propulsion engineers, Stoke proposes a fully reusable two-stage rocket using a ring-shaped aerospike engine; neither rocket nor engine had flown, making vertical integration possible but plainly risky.
The proposed economics did not yet convince Diamandis. Launch might fall toward $100 per kilogram from $500–$1,000, but space power at $6–$9 per watt versus $12 terrestrially was “not enough of a difference for the level of complexity”; he wanted nearer a 10-fold advantage. Compact fusion and heat dissipation remained debated.
China’s CosmoSpace was said to be planning three orbital modules: one at 100-megawatt power, one with 10 terabits per second of communications and one delivering 10 exaflops. Wissner-Gross called it part of a race toward Dyson swarms; the entire orbital-data-center conversation had moved from fringe to weekly topic in roughly four months.
Mostaque’s security objection was blunt: orbital systems could “just start disappearing.” Blundin answered that three-year hardware depreciation only requires protection through payback, perhaps under Space Force coverage. Solar activity already caused errors in Mostaque’s terrestrial A100 training, so space systems need granular fault tolerance, checkpointing and recovery rather than cluster-wide rollbacks.
19. Humanoid power is advancing faster than its rules of engagement
A possible 2026 U.S. robotics executive order could bring tax credits, subsidies and trade protection analogous to semiconductor policy. The competitive prompt is China, which installed 54% of the world’s robots in the cited year and has more than 150 Chinese robot companies despite official concern about a bubble.
Wissner-Gross called humanoids “the next big thing for AI after agents,” capable of reindustrializing the West and automating the two-thirds of services requiring physical intervention. Blundin called for robot athletics to support Western humanoid development and make performance visible.
Optimus and Figure appeared markedly more natural while running than six months earlier, but Engine AI’s 5-foot-8 T800 shifted the conversation from locomotion to force. Mostaque cited maximum joint torque of 450 N·m and claimed it could punch harder than a gorilla or four times Mike Tyson.
Diamandis objected that copying the human form ignores the wheel’s energy efficiency and proposed adding wheels. Mostaque raised the governance question of whether street humanoids need torque limits. Dave raised warfare as the more frightening application, while Salim described governance of different rules of engagement for kitchens, streets and battlefields.