Claude Opus 4.5, Genesis Mission and Amazon's $50B AI Push
Summary
Genesis Mission recasts U.S. basic science as national AI infrastructure, combining Department of Energy supercomputers, federal datasets, and laboratory tools to compress research from years to days. Alexander Wissner-Gross called it a “1939 moment” in which America becomes “one big AI factory”; Peter Diamandis highlighted biotech, fusion, and quantum, while the stated goal is to double U.S. scientific productivity over a decade. The upside depends on proper funding and execution. Wissner-Gross’s key condition: open science could make the impact “truly exponential,” whereas closed public-private work with strong IP protections would have much lower impact.
Claude Opus 4.5 is presented as a possible threshold for recursive self-improvement, not merely another benchmark release. The headline claims were 76% fewer tokens for equivalent results, leadership in seven of eight programming languages, and AI performance exceeding incoming Anthropic performance-team employees on key assignments—the “canary” for Wissner-Gross. Mostaque reported 52% on SWE-bench Pro without reasoning tokens versus 45% for his Intelligent Internet framework, a 67% cost reduction to $25 per million tokens, and a path to one-shotting typical 100,000–200,000-token codebases next year.
The cost of capable intelligence is falling as evaluation shifts from puzzle scores toward dollars earned. Opus reportedly scored 75% in a same-model multi-agent setup and 88% when orchestrating Haiku or Sonnet, while Wissner-Gross led with, “We’re driving the cost of intelligence to zero.” Mostaque expects benchmarks such as Vending Bench and trading tests to measure real economic output next; he would be surprised if a single entrepreneur could not build a billion-dollar business within two years, “probably next year,” while Wissner-Gross argued an altcoin-pumping “baby AGI” could do it now-ish.
AI-native companies could invert the labor-and-capital model by making nearly every operating input variable cost. Mostaque’s mechanism is unusually concrete: enterprises can charge customers upfront, pay AI providers one or two months later, and automate compliance, forecasting, tax, and payments—potentially allowing a complete business to launch “in minutes” in about a year. Salim Ismail connected this to near-zero acquisition and supply costs, while Wissner-Gross argued agents are “neither capital nor labor” and humans may become investors in fleets of AI entrepreneurs.
The compute trade is broadening from Nvidia scarcity into a heterogeneous market spanning Google TPUs, AWS Trainium, memory, power, and interconnects. Google’s seventh-generation Ironwood TPU was described as four times faster than its predecessor; Mostaque emphasized Google’s chip interconnects, million-to-2-million-token context, and an environment where DRAM prices had risen about fivefold. Amazon, meanwhile, plans up to $50 billion of U.S.-government AI infrastructure and 1.3 GW of new capacity starting in 2026, while its $11 billion, 2.2-GW Indiana facility runs 500,000 Trainium2 chips largely suited to inference.
Shopping agents turn control of user intent into the next distribution battle. ChatGPT’s shopping research, using ChatGPT Mini, claimed up to 64% accuracy, but Mostaque contrasted that with Amazon Rufus’s reported 250 million users, conversion rates up to 60% higher, and an estimated $10 billion of incremental sales next year. The winning layer may be the secure, charming “Jarvis” beside the user—observing requests, conversations, and eventually gaze—then routing work across specialist agents while disintermediating search, affiliate media, and recommendation engines.
The panel sees severe labor disruption arriving before abundance, making coordination and economic growth the binding policy problems. Mostaque’s forecast is that most keyboard-and-mouse work becomes “negative value” within at most 900 days, though he explicitly did not predict every job disappearing; he also cited Grok 4.1 Fast scoring roughly 95% on TaoBench at $0.50 per million words and predicted no customer-service jobs within two years. Proposed bridges included universal AI, AI social scientists, UBI, universal basic services, and universal basic equity—but Wissner-Gross’s prerequisite was to grow the economy faster than conventional human labor loses value.
Brain-computer interfaces and falling launch costs remain the episode’s high-upside physical-world bets. Paradromics was said to reach 200 bits per second versus Neuralink’s roughly 10, with approval to begin human testing in about two months; Ismail, formerly a “hard no” on high-bandwidth BCI, conceded, “Oh, shit. He’s right again.” Diamandis separately traced launch costs from roughly $50,000/kg for the shuttle to $2,500 for Falcon 9, a projected $100 for Starship, and potentially $0.10/kg for lunar mass drivers—cost curves that would expand usable land and material supply far beyond Earth.
Deep dive
1. Genesis turns U.S. science into one compute system
Diamandis framed the executive order as a unified platform connecting Department of Energy supercomputers, laboratory data, and previously isolated federal datasets. Its intended payoff is to shrink AI-driven experimentation from years to days across biotech, fusion, and quantum research; he called it potentially the greatest U.S. accelerator for human knowledge if properly funded and executed.
Wissner-Gross’s analogy was deliberately martial: “This is the Manhattan Project.” Where 1939 turned America into one factory for nuclear weapons, Genesis would turn it into “one big AI factory,” supplying both compute and the “limiting reagents”—datasets, software tools, and scientific infrastructure—that models need.
Mostaque argued the Department of Energy’s leadership is no coincidence because energy is itself the constraint. More deregulation, fusion, and solar could reinforce the compute program; solving even one major energy or science problem could have immense consequences, assuming Genesis is properly executed.
Ismail called the program both catch-up—China and France have used sovereign datasets for years—and an example of government at its best. His old DoD lesson was that venture capital waits near the exponential curve’s knee, but government funds “the arbitrarily long flat part”; Genesis could pull that knee forward.
Wissner-Gross said success also turns on openness: a Manhattan-Project-style closed effort could limit spillovers, while open science would be “truly exponential.” Public-private partnerships with strong IP protections and substantial private work would have much lower impact.
2. Opus 4.5 is the recursive-improvement canary
The release arrived with unusually specific claims: 76% fewer tokens for equivalent results, leadership in seven of eight programming languages, and a 15% improvement in multi-agent support. Wissner-Gross cared most that the model reportedly outperformed incoming Anthropic performance-team employees on key assignments.
His definition of recursive self-improvement is operational: frontier labs begin allocating more compute and infrastructure to AI researchers than human researchers, allowing models to research and code better descendants. “We’re nearing, if not already at the point,” he said, with the caveat that Opus’s pretraining cutoff was months earlier.
Mostaque’s own comparison sharpened the claim. Intelligent Internet had achieved 45% on Scale AI’s SWE-bench Pro using multiple models; Opus 4.5 reportedly reached 52% without reasoning tokens. That surprised him because recent gains had depended on models thinking longer, checking work, and spending more inference compute.
Opus alone reportedly scored 75% in a same-agent multi-agent test, rising to 88% when paired with cheaper Haiku or Sonnet agents. Mostaque called it the first AI that could “provably” supervise other agents, opening the swarm model; Wissner-Gross added that his one-shot Mario-style side-scroller test also worked beautifully.
3. Benchmarks are shifting from puzzles to dollars
ARC-AGI 1 and 2 test visual pattern recognition and program synthesis on tasks humans find intuitive but models historically found difficult. Wissner-Gross cited Opus’s cost efficiency and a result in which a company named Poetic announced superhuman-level ARC-AGI 2 performance as evidence that these benchmarks are beginning to saturate.
The economic consequence extends beyond visual puzzles: physical-world automation requires recognizing patterns, manipulating environments, and synthesizing implicit procedures. “The world needs harder benchmarks,” Wissner-Gross concluded, because capabilities once separating humans from AI are being solved.
Mostaque expects the next evaluation axis to be literal dollars—Vending Bench, trading results, and other tests of sustained economic work. Models are moving from “very smart people you tap on the shoulder” for isolated tasks to agents capable of operating businesses and earning money.
4. AI agents turn companies into variable-cost portfolios
On the billion-dollar solo company, Mostaque’s estimate was “a year or two away at most”; Wissner-Gross said now-ish, albeit perhaps first through an AI pumping an altcoin. Ismail noted that one acquaintance had already launched 47 AI startups in a month: the production engine exists, but one venture still needs product-market fit.
Wissner-Gross rejected the usual capital-versus-labor framing because agents may be “a new third category.” His alternative to everyone becoming an entrepreneur is everyone becoming an investor, owning fleets, indices, or entire economies of AI agents that identify and solve valuable problems.
Mostaque explained why the company stack changes: hiring, servers, compliance, forecasting, tax, and payment reconciliation become on-demand services. Because customers can pay upfront while enterprise AI bills settle later, a new company can be both variable-cost and cash-flow-positive; he expects the full launch stack to be ready in about a year.
Ismail connected this to zero marginal cost on both sides of the business. The internet collapsed demand acquisition costs; Airbnb-like models collapsed supply costs; AWS removed computing from the balance sheet. Remove both constraints and “the market cap explodes”—the organizational “holy grail.”
5. Shopping agents make distribution and trust the choke point
Diamandis read ChatGPT shopping research as an attack on search engines, affiliate blogs, YouTube reviewers, and Amazon recommendations. The system reportedly uses ChatGPT Mini and reaches up to 64% accuracy in selecting products aligned with what a user wants.
Wissner-Gross saw a broader rejection of the single all-purpose model. Alongside general models, OpenAI had launched Deep Research, Codex, and now a shopping specialist; finance, medicine, and consulting could follow. Routing can still hide that proliferation behind “a single pane of glass.”
Ismail asked when a private, sovereign Jarvis would choose agents and act for the user. Wissner-Gross answered, “Now,” pointing to computer-use agents already deployed or entering beta; heads-up displays and wearable interfaces are the next surface, not a prerequisite for the underlying capability.
Mostaque argued distribution may dominate raw model quality. He recalled Amazon CEO Andy Jassy saying Rufus had 250 million users, with conversion statistics reportedly up to 60% higher and potentially $10 billion in sales next year. Routine purchases can disappear into automation; discretionary shopping may belong to whichever agent is most trusted, charming, and ever-present.
6. Google makes accelerated compute a multi-supplier market
Google’s seventh-generation Ironwood TPU was presented as offering four times its predecessor’s performance, with TPUs increasingly available as cloud capacity and potentially inside customers’ own data centers. Meta’s reported use illustrates direct competition with Nvidia without requiring customers to purchase the hardware.
Wissner-Gross welcomed the erosion of the alleged Nvidia/CUDA monopoly. Google TPUs, AMD accelerators, Amazon Trainium, and specialized ASICs are turning accelerated compute into a “multi-supplier, very healthy, very heterogeneous ecosystem” rather than a single-vendor bottleneck.
Mostaque emphasized Google’s advantage in joining chips: from 64 per unit previously to what he believed was roughly 9,000, with prior runs reaching 50,000 lower-energy chips. He said the current generation delivers about 10 times the compute of the v5s, and Google can link systems across data centers.
With DRAM prices up about fivefold, his key distinction was RAM versus FLOPs. Google’s architecture supports Gemini inputs of 1 million to 2 million tokens across video, audio, and text; as raw model speed becomes adequate, Mostaque expects context capacity to become the more important performance differentiator.
7. Amazon is paving farmland with inference capacity
Amazon plans up to $50 billion for U.S.-government AI infrastructure, adding 1.3 GW of capacity with construction beginning in 2026. Wissner-Gross called government clouds notoriously GPU-starved despite the public sector representing roughly one-quarter to one-half of the economy, depending on measurement.
Diamandis cited AWS serving 11,000 government agencies and expecting $125 billion of capital expenditure by the end of 2025. Ismail asked whether AWS had effectively become the federal government’s cloud; Wissner-Gross pushed back that government uses multiple providers, though AWS is plainly a key supplier.
Project Rainier in rural Indiana compresses the industrial story into one site: 1,200 acres, seven buildings erected in roughly a year, $11 billion invested, 2.2 GW of power, and 500,000 Trainium2 chips. “We’re tiling the Earth with compute,” Wissner-Gross said, comparing converted Midwestern farmland to 1939 factory mobilization.
Mostaque characterized Trainium2 as Hopper-equivalent and solid for Claude inference, though about one generation behind for demanding training runs. Training requires resilient interconnects and backpropagation across many devices; inference is mostly forward matrix multiplication, making it easier to port through frameworks such as OpenXLA.
8. Falling launch costs expand the map of scarcity
Diamandis’s launch-cost ladder began with the shuttle: planned at $50 million and 50 flights annually, but ultimately roughly $1 billion–$2 billion per launch, one to four flights a year, and about $50,000/kg. Falcon 9 cut that to approximately $2,500/kg through first-stage reuse.
His projected next steps were another 25-fold decline to $100/kg with fully reusable Starship, then lunar mass drivers at roughly $0.10/kg using electromagnetic acceleration and solar electricity. Ismail stressed that this is a logarithmic cost collapse in “a very physical environment,” effectively making launches increasingly software-like.
The disagreement was over destination and resource use. Diamandis preferred the Earth-Moon system, asteroid-built O’Neill cylinders, and avoiding another planetary gravity well; Wissner-Gross cheerfully advocated “disassembling the solar system,” while conceding the asteroid belt could serve as “training wheels.”
Mostaque’s Earthbound counter to land scarcity was 16 billion habitable acres—about two per person—potentially rising to 20 billion–22 billion as passenger drones and cheap energy, heating, and cooling open difficult terrain. He projected population peaking around 10 billion by 2050; Ismail added a cited Tesla robotaxi forecast of about $0.30/mile, which could radically expand practical commuting geography.
9. BCI bandwidth is rising before neuroscience is understood
Paradromics had completed sheep testing and was approved to begin human testing in roughly two months, around January or February. Its claimed 200 bits per second is 20 times Neuralink’s roughly 10; Ismail, once a “hard no” on high-bandwidth BCI in the early 2030s, admitted the progress had changed his mind.
Mostaque expects BCI to become one of the largest investment areas over the next three years because it can both treat impairment and augment humans. Ismail pointed to Max Hodak, Neuralink’s co-founder, and his company Science Go, which uses neural stem cells, as well as Merge Labs, in which Sam Altman invested, as evidence that architectures are proliferating.
Wissner-Gross’s speculative curve puts a gigaflop of compute at the size of a human brain cell around 2045. He separately expects a full high-resolution human connectome probably within five years, while acknowledging that today’s destructive slicing is extremely invasive and that a connectome is only a proxy for an upload.
Mostaque argued useful decoding may arrive without nanotechnology: Stability’s Mind’s Eye reconstructed viewed images from low-bandwidth MRI, and diffusion could infer rich brain processes from partial signals. Wissner-Gross placed Drexlerian assemblers no later than 2045 but considered softer DNA-origami or AI-designed versions plausible within 10 years.
10. The abundance transition is a coordination problem
Mostaque restated his forecast carefully: within at most 900 days, most human economic work performed through a keyboard or mouse becomes “negative value.” He did not say every job disappears. Wissner-Gross replied that, if AGI means generality, it has existed since at the latest summer 2020 with GPT-3 and the “Language Models are Few-Shot Learners” paper, while economically general AI is either here or likely imminent, depending on the benchmark.
The sharper labor warning came from Grok 4.1 Fast, which Mostaque said scored roughly 95% on TaoBench at $0.50 per million words. “No customer service jobs within two years,” he predicted, adding that adoption takes time but the direction is one-way.
Mostaque’s remedy combines top-down AI social scientists, such as the Sage project, with universal personal AI that makes poor and otherwise “invisible” people legible to food, healthcare, and benefit systems. Governments should coordinate with AI, distribute AI directly, and study responses such as the 1933 New Deal.
Ismail proposed UBI or universal basic services during a difficult 10-year shift from scarcity institutions to abundance, while Wissner-Gross added universal basic equity. His governing test was macroeconomic: grow total output faster than conventional labor is obsoleted, because “it’s relatively easy to distribute abundance if we have abundance.”
11. Their 2035 horizon assumes the hardest curves keep compounding
Ismail imagined Thanksgiving dinner costing one-tenth as much, personalized to each metabolism and produced with ultra-cheap food and energy. Wissner-Gross wanted some humans celebrating on Mars or in the cloud, perhaps alongside uplifted animals; Mostaque called 10 years “the pessimistic end of the AGI forecast,” conditional on humanity navigating the transition.
Wissner-Gross’s defining 2025 marker was that mathematics is “credibly and definitively being solved by AI,” a canary for grand challenges in science, engineering, and medicine. Diamandis chose humanoid robotics and the capital now flowing into manufacturing as evidence that his own “Data or C-3PO” is approaching.
Mostaque was grateful that AI social scientists and the required coordination infrastructure now look buildable—and that, “minus the hard light,” the tools for a holodeck exist. Diamandis closed on longevity escape velocity, citing hyperscaler interest and Dario’s stated ambition to double human lifespan within five to 10 years.