Pioneers Insight Method Research Author
Google I/O 2026, Karpathy Joins Anthropic, and Cerebras’ $95B IPO | EP #256
Back to Episodes

Google I/O 2026, Karpathy Joins Anthropic, and Cerebras’ $95B IPO | EP #256

Summary

  • Google’s comeback is a distribution-and-infrastructure thesis, not a single-model victory. Sundar Pichai said monthly processing rose from 9.7 trillion tokens two years ago to 480 trillion last year and 3.2 quadrillion today, while Gemini reached 900 million monthly users and AI Mode passed 1 billion. CapEx climbed from $31 billion in 2022 to a projected $180–190 billion this year, yet the stock rose—the counterfactual Dave Blundin said nobody would have believed five years ago—and Google is now “disrupting the disruptors.”

  • Google’s technical portfolio is bifurcating between a distinctive multimodal bet and throughput-led models the panel did not mistake for intelligence leadership. Gemini Omni can generate and conversationally edit video from text, photos, video, and audio; Alexander Wissner-Gross called multimodality Google DeepMind’s increasingly “idiosyncratic bet” as US rivals focus elsewhere. Gemini 3.5 Flash, meanwhile, claims four-times-faster output than other frontier models, but Wissner-Gross rated its raw capability “solidly mid” and saw the release as throughput- and tool-use-maxing ahead of Gemini 3.5 Pro.

  • Antigravity 2.0 and Gemini Spark may be fast-following products, but Google’s defaults make them strategically dangerous. Antigravity adopts the agent-first interface already associated with Cursor, while Spark runs continuously in a Google Cloud VM and acts across Gmail, Sheets, and other services. Wissner-Gross’s distribution argument was blunt: Google can be “as good one day later,” integrate with Search, Chrome, and Android, and make billions of users’ first persistent agent a Google agent.

  • Search is becoming a persistent intent-and-commerce layer capable of steering both questions and transactions. Google’s agents can continuously scan for apartments or price changes, while AI suggestions can reshape the query itself; Blundin flagged the “astounding” monetization power for a business protecting roughly $200 billion of 90%-margin revenue. Universal Cart then connects Search, Gemini, YouTube, and Gmail to deal monitoring and checkout—the path Salim Ismail summarized as “intent to agent to transaction,” with the deeper disruption being “getting rid of shopping.”

  • Synthetic abundance makes verification, rather than creation, the scarce asset. SynthID has watermarked more than 100 billion images and videos plus 60,000 years of audio, and OpenAI, Kakao, and ElevenLabs are joining NVIDIA in adopting it. The panel framed this as industry self-regulation moving faster than government and as a transition “from the information age to the verification age,” where provenance becomes infrastructure.

  • Google is extending AI simultaneously into ambient interfaces, scientific research, and rapid entrepreneurship. Its first audio-only intelligent glasses arrive in the fall, trading a visual display for battery life and private spoken assistance while intensifying concerns about recording, backlash, and “the loss of presence in life.” Gemini for Science targets papers, code, hypotheses, weather simulation, and drug discovery; the Build with Gemini XPRIZE gives teams 90 days to address a problem affecting at least 100,000 people, though the presentation cited $2 million while Peter Diamandis later described $3 million in prizes plus $1 million for operations.

  • Andrej Karpathy’s move to Anthropic signals that frontier knowledge and compute are concentrating inside a handful of labs. He will use Claude to accelerate Claude’s own pre-training research after warning that, outside a frontier lab, one’s “judgment fundamentally will start to drift.” The panel read his choice of Anthropic—and his willingness to re-enter a large institution after OpenAI, Tesla, and Eureka Labs—as evidence that independent researchers risk missing the formative inner loop.

  • Cerebras’ 68% IPO surge to a stated $95 billion market cap validated a wafer-scale inference bet that spent years ahead of demand. Andrew Feldman said Cerebras built a dinner-plate chip 58 times larger than any predecessor, filled it with SRAM, and now runs inference 15–20 times faster than GPUs; system sales progressed from 12 in generation one to 300–350 in generation two and “many, many thousands” in generation three. The company subsequently signed a more-than-$20 billion multiyear OpenAI deal, but Feldman warned that domestic fabs remain a decades-long constraint and put Elon Musk’s Terafab ambition in the 15–20-year category.

Deep dive

1. Google’s sixfold CapEx jump bought a credible full-stack comeback

  • Pichai’s scale-setting numbers ran from 9.7 trillion monthly tokens two years ago to 480 trillion last year and 3.2 quadrillion today. Google now reports 8.5 million monthly model builders, APIs processing 19 billion tokens per minute, 13 products above 1 billion users, and five above 3 billion.

  • Consumer reach is following: AI Overviews has 2.5 billion monthly users, AI Mode has already exceeded 1 billion, and Gemini rose from 400 million to more than 900 million in a year. Nano Banana models have generated over 50 billion images.

  • Infrastructure spending rose from $31 billion in 2022 to a projected $180–190 billion this year. Google can distribute training across more than 1 million TPUs globally, while its new chips claim up to twice the performance per watt—making the stack competitive “from the transistor all the way through the user experience.”

  • Blundin’s investor counterfactual: nobody five years ago would have believed Google could sixfold CapEx while its stock appreciated. Wissner-Gross called the outcome partly inevitable: Google had the Transformer, compute, and an exponentially lengthening search interaction, but had to assemble the institution before it could start “disrupting the disruptors.”

2. Omni makes modality scaling Google’s clearest differentiated wager

  • Gemini Omni is pitched as “create anything from any input”: text, photos, video, and audio can feed realistic clips, interactive simulations, or conversational edits. Demis Hassabis demonstrated a claymation protein-folding explainer that accurately rendered amino-acid chains, alpha helices, beta sheets, and three-dimensional structure.

  • Blundin said Hassabis received the event’s strongest audience response because medicine, science, and playful self-transformation made the capability tangible. Diamandis emphasized the educational potential, while Diamandis and Ismail discussed generating graphics and visual or audio tracks in real time alongside live conversation.

  • Wissner-Gross called Google DeepMind arguably the only American frontier lab still seriously pursuing video and broad multimodality, contrasting it with OpenAI’s Sora retrenchment and Anthropic’s CodeGen focus. His larger inference was that Google may treat DNA, protein sequences, and “dozens of other modalities” uniformly—a distinctive but still unproven route toward superintelligence.

3. Gemini 3.5 Flash wins on speed, not the panel’s frontier crown

  • Google made Gemini 3.5 Flash the default model in the Gemini app and AI Search Mode, claiming it beats 3.1 Pro on almost every benchmark, including a large GDPval jump. Its headline positioning is frontier-comparable intelligence with output four times faster than other frontier models.

  • Wissner-Gross’s dissent was direct: “I would call Gemini 3.5 Flash solidly mid.” On raw capability, he said it compared poorly with GPT-5.5 High, X High, or Pro, while acknowledging that Flash is not the top-range release and Gemini 3.5 Pro was promised for another month.

  • His benchmark reading was that Google chose the most flattering intelligence-versus-throughput axis and emphasized tests rewarding aggressive tool use. He placed Google at “tier one and a half,” interpreting Flash as throughput-maxing and tool-use-maxing for latency-sensitive internal products rather than a new raw-intelligence frontier.

  • Blundin relayed Silicon Valley’s “two-horse race” view favoring OpenAI and Anthropic for the smartest enterprise AI, but noted Google’s $180 billion CapEx base is no longer uniquely overwhelming after OpenAI raised a stated $120 billion in cash. Ismail’s synthesis was a durable bifurcation between premium cognition and ultra-cheap, fast cognition as “marginal intelligence cost trends towards zero.”

4. SynthID turns synthetic provenance into a trust layer

  • SynthID has watermarked over 100 billion images and videos plus 60,000 years of audio. Google is adding credentials that distinguish camera capture from AI generation or editing; OpenAI, Kakao, and ElevenLabs are joining NVIDIA as adopters, although scale still depends on broader participation.

  • Wissner-Gross expects “cryptographic chains of custody” analogous to a browser’s lock icon, calling the destination a “proof of reality.” The irony is that authentication may begin at the synthetic end—vendors claiming provenance for generated content—before camera and recording manufacturers join the same protocol.

  • Diamandis framed watermarking as an early act of industry self-governance because lawmaking cannot match AI’s pace. Ismail’s formulation was “scarcity equals abundance minus trust”: as creation costs collapse, value moves toward filtering and authenticity, taking society “from the information age to the verification age.”

5. Antigravity 2.0 assumes developers stop touching code

  • Antigravity 2.0 is an “unabashedly agent first” desktop application for parallel agent conversations and artifacts. Google’s demo asked it to repair an operating system that could not run Doom; it researched the missing video and keyboard drivers, wrote more than 100 lines, rebuilt the OS, and launched the game.

  • Wissner-Gross traced the product to Google’s Windsurf acquisition and called the new interface a fast-following version of Cursor’s shift from a VS Code-style editor to agent orchestration. His experience with Antigravity 1.0 was buggy, and he knew no primary developers choosing it over Claude Code, Codex, or Cursor.

  • Blundin’s hands-on test found 2.0 almost indistinguishable from Cursor, except Google went further by hiding code entirely and requiring users to launch the old interface to edit it. That makes functional evaluation—“I don’t like where that button is, move it”—the new development layer.

  • Wissner-Gross’s arrow of time was categorical: “Code is clearly going away as a human endeavor,” followed by recursive model self-improvement. Blundin expects the next interface to become a real-time, graphical “Star Trek holodeck,” with objects manipulated directly rather than regenerated between prompts, potentially within the calendar year.

6. Gemini Spark makes Google’s bundle an agent-distribution moat

  • Gemini Spark is a continuously running agent hosted on dedicated Google Cloud virtual machines and powered by Gemini 3.5 Flash. It can compile launches into an email, generate an RSVP tracker, and update the Sheet automatically when a reply arrives in Gmail—persistence married to cross-product access.

  • Wissner-Gross called it a “lazy copycat product” and the minimum viable response to OpenClaw: host a VM, connect Google products, and stop short of inventing next-generation agent benchmarks or capabilities. Diamandis also preferred the explicit personality of his OpenClaw-based “Skippy” to a generic assistant.

  • Diamandis saw historical irony in Google tying products together after antitrust action stopped Microsoft from using Windows and Internet Explorer to foreclose challengers such as Google; he recalled the case ending with a $1 fine and forced unbundling. Google is now exploiting Search, Gmail, Chrome, Android, and billion-user products in much the same strategic shape.

  • Ismail conceded Spark is “boring” and “safe” but argued that hundreds of millions could learn agent behavior through it. Wissner-Gross would not underestimate defaults: Google does not need the frontier risk if it can be “as good one day later,” remove OpenClaw’s installation friction, and place the first personalized agent one click from Search.

7. AI Search preserves the rectangle by making it persistent

  • Google’s reimagined search box expands as a user talks, suggesting nuances beyond autocomplete and managing multiple agents. An apartment agent can absorb a “total brain dump” of criteria, then continuously scan websites, social platforms, and forums for newly available properties.

  • Diamandis worried that assistance could become direction: a Barbados query might be nudged toward Bermuda. Blundin called that steering power “astounding” for protecting roughly $200 billion of existing 90%-margin revenue; rather than placing obvious ads beside answers, Google could influence the question and default recommendation.

  • His TripAdvisor analogy preserved the mechanism: reviews could remain accurate while paying hotels were sorted upward, because roughly 80% of users clicked among the first two or three choices. People are not merely lazy; decision overload forces them into defaults, which lets Google monetize ordering without recommending fraudulent products.

  • Wissner-Gross’s symbolic takeaway was that “the shape of the rectangle changed after decades”: Google self-disrupted the familiar search box before ChatGPT could own the new interaction. A later guest added that the historical obstacle may have been technical—generative models were too slow and costly for search—which explains Flash’s unusually strong throughput emphasis.

8. Universal Cart shifts commerce from browsing to delegated intent

  • Universal Cart spans Search, Gemini, YouTube, and Gmail across merchants including Nike, Target, Walmart, and Shopify. Once an item is added, the cart monitors deals, price history, price drops, and restocking, allowing the transaction to remain active after the user stops shopping.

  • Wissner-Gross identified Amazon as “the elephant in the room”: Google keeps building virtual storefronts and standards to challenge Amazon, but Amazon may simply decline to participate. Blundin nevertheless viewed consumer retail as secondary to the larger AWS-versus-Google Cloud battle already consuming the companies.

  • Ismail compressed the new funnel from human–website–cart–checkout to “intent to agent to transaction.” Marketers may need to persuade 100 million purchasing agents, but Diamandis and Ismail pushed further: an agent that knows taste and context can send returnable surprises, meaning “the disruption to e-commerce is not the better shopping, it’s getting rid of shopping.”

9. Gemini’s integration story is weakened by Google’s product sprawl

  • The Gemini app now reaches more than 900 million users, while NotebookLM has produced over 1.5 billion notebooks, podcasts, slide decks, and other artifacts. It is available in more than 230 countries and over 70 languages, with regional dialect selection coming to its audio output.

  • Product demonstrations ranged from transforming a musician’s raw footage and reference images into a promotional video to Daily Brief, an agent synthesizing email and travel information with actions embedded inline. Blundin said the crowd responded more strongly to the recognizable musician than to an agent building an operating system.

  • The panel’s pushback was organizational: why do NotebookLM, Spark, Flash, and Antigravity remain separate brands when AI can unify them behind one interface? Wissner-Gross asked for “more wood behind fewer arrows” and invoked the abandoned Google Reader; Ismail compared the resource dilution to Yahoo’s “peanut butter problem,” where launching earns rewards but sustained iteration does not.

10. Audio glasses trade visual ambition for battery life—and raise a social bill

  • Google’s first Android XR audio glasses are scheduled for the fall, with Samsung and eyewear partners—Orby Parker and Gentle Monster. Forward-facing cameras feed Gemini, while responses are privately spoken; users can take photos, call, play music, and access phone apps without reaching into a pocket.

  • Ismail found the absence of a display “very mid,” though he saw continuous human-computer interaction becoming an ambient layer. Ismail and Diamandis criticized Google’s lapse: Google owned the early Glass opportunity, then abandoned consumer iteration while Meta spent billions and established a lead that Google and Apple now chase.

  • Blundin warned that the product could split society between people recording continuously and people offended by being recorded. The old risk of getting “punched in the face” is not comic nostalgia: Wissner-Gross cited three commencement events where crowds booed as soon as speakers mentioned AI, evidence that Silicon Valley’s enthusiasm is not socially representative.

  • Wissner-Gross defended audio as a practical intermediate design: no display means better battery life, and visual material can still appear on a paired phone. Yet the discussion returned to attention—an agent whispering identities, purchases, and alerts may be useful, but “the loss of presence in life can be really costful.”

11. Gemini for Science and XPrize aim the stack at real problems

  • Gemini for Science combines literature monitoring, conversion of research goals into usable code, hypothesis generation, and simulation. WeatherNext is intended to predict hurricane paths faster and more accurately than traditional methods, while Isomorphic Labs is in preclinical work on multiple projects for immune disorders and cancer.

  • Wissner-Gross endorsed DeepMind tackling what Hassabis calls “root node problems,” such as protein folding and fusion. He saw the scientific tools as meta-science—algorithms for producing more science—and potentially more public-benefit work than a deeply monetized Google business.

  • The Build with Gemini XPRIZE announcement contained two stated funding figures: the onstage clip offered $2 million, while Diamandis later said Google supplied $3 million for prizes plus $1 million for operations. Entrants must choose a problem affecting at least 100,000 people and build, market, and generate revenue within 90 days.

  • Diamandis framed the contest as teaching people to fish: describe the product, market, and interfaces in English, then let agents build them. Ismail’s specimen was a gamified app that substantially addressed lazy eye, illustrating how open innovation can make an obscure but widespread problem actionable; finalists are scheduled to pitch on September 25.

12. Karpathy chose frontier access over independent distance

  • Andrej Karpathy joined Anthropic’s pre-training team to start an initiative using Claude to accelerate Claude’s own pre-training research. The episode traced his path from co-founding OpenAI to leaving in 2017 for Tesla full self-driving, returning to OpenAI in 2023 and 2024, and founding Eureka Labs.

  • Before the announcement, Karpathy said outsiders cannot see what frontier laboratories have coming: their “judgment fundamentally will start to drift,” the systems remain opaque, and they lose an under-the-hood understanding. He floated spending time inside a lab, doing serious work, and then perhaps returning to independence.

  • Blundin interpreted the decision as a compute constraint: Karpathy’s auto-research repository wanted far more compute than an independent researcher could command, and “you can’t miss” the singularity outside the large machines. Feldman generalized the lesson to hardware—without fundamental engagement with Google, Anthropic, or OpenAI, chip designs can drift from frontier requirements too.

13. The Musk–OpenAI verdict was treated as a costly distraction

  • A federal jury unanimously rejected Elon Musk’s OpenAI lawsuit after two hours of deliberation, finding his claims outside the statute of limitations; his legal team planned an appeal. Andrew Feldman said the appeal would very likely fail because the ruling rested on factual timing, while noting that the same issue could seemingly have ended the case earlier.

  • Feldman dismissed “what billionaires in pissing matches” do as uninteresting. He praised Musk as a breathtaking polymath and Sam Altman as the builder of one of capitalism’s fastest-growing companies, but said “everybody loses when they battle”—he wants both spending their time building rather than litigating.

14. Cerebras spent years early, then inference demand caught up

  • Cerebras’ IPO raised a stated $5.5 billion, closed 68% higher, and produced a $95 billion market capitalization; the hosts described it as the largest US IPO since Uber in 2019 and the third-largest technology IPO. Feldman made the listing a family event and called sharing it with employees, parents, and family “really something special.”

  • The founders began meeting in 2015 after AMD acquired their previous startup in 2012. Their two contrarian bets were that AI would become large enough to require dedicated silicon and that winning required a “clean sheet of paper,” not another GPU derivative.

  • Cerebras pursued memory bandwidth through a dinner-plate chip 58 times larger than any chip previously built, “stuffed to the gills with SRAM.” The fast but area-intensive memory became viable at wafer scale, producing what Feldman described as 15–20-times-faster inference than GPUs.

  • Solving wafer scale in August 2019 did not create immediate demand: generation one sold 12 systems, generation two 300–350, and generation three “many, many thousands.” Useful inference finally arrived in late 2024 and early 2025; Cerebras later signed an OpenAI agreement above $20 billion over several years and a March term sheet for AWS deployment.

15. Wafer-scale SRAM changes the relevant inference metric

  • Feldman readily admitted Cerebras “got an enormous amount wrong,” but correctly separated how AI is made—training—from how it is used—inference. Between 2020 and 2024–25, models were not useful enough; once customers stopped asking about parameter counts and started asking whether models wrote good code or completed work, inference demand overwhelmed the company.

  • The third-generation wafer-scale engine has roughly 40–50 gigabytes of SRAM, so trillion-parameter models still span chips. Cerebras partitions them so no layer crosses more than two chips, then moves the relatively small result vector over 100-gigabit Ethernet; Feldman put the resulting penalty near 2%.

  • By contrast, he said an 800-square-millimeter SRAM design such as Groq’s—which he said NVIDIA acquired—might divide a large model across 2,000–3,000 chips and pay for every hop. Cerebras published roughly 1,000 tokens per second on “Kimiko 2” versus 70 at Fireworks: about a 15-times advantage.

  • Feldman cautioned that “tokens per second” can conceal aggregate throughput. An NVL72 might generate millions of slow tokens at 35 tokens per second, yet support only one or two users at 200 tokens per second each despite costing $4 million; the investable metric is per-user latency and concurrent service, not a rack’s gross number.

16. Fabs and packaging—not just chip designs—set the supply ceiling

  • Asked about Musk’s Terafab plan to produce 50 times today’s global chips, Feldman separated vision from manufacturing reality. He called Musk uniquely capable but said fabs “will always take longer than he says” and cost vastly more, placing the undertaking at 15–20 years rather than five or ten.

  • Even experienced builders need five or six years and $40–50 billion for a fab. Identical ASML equipment does not erase accumulated process knowledge: TSMC remains ahead of Samsung because generations of received wisdom matter, and US projects must survive changing administrations, local ordinances, and unforeseen construction problems.

  • Reshoring fabrication alone is insufficient because packaging “breathes power and life” into dead silicon through power and I/O. The US also surrendered deposition, materials, process, and packaging expertise; much now sits in Taiwan and Korea, with important materials supplied from Japan by companies such as Kyocera.

  • Cerebras has committed its 3-nanometer design to TSMC, manufactures some components through Samsung, and has never used Intel. Feldman respects Intel leader Lip-Bu Tan but said substantial work remains before Cerebras could move there; a finished wafer meanwhile passes through numerous partners for backside layers, dicing, cleaning, packaging, and power delivery.

17. Cerebras’ moat was forged through expensive failure before rivals met the same problems

  • The first wafer-scale chip took roughly four years and $400–500 million, spanning lithography, architecture, packaging, cooling, power delivery, compilers, and algorithms. Feldman jokes that he takes the wafer to dinner because it embodies a pioneering object no one had successfully built in the computer industry’s 75-year history.

  • That early work exposed packaging failures years before competitors encountered them. Feldman said the “B100—or the B200” arrived 18 months late because of CoWoS and coefficient-of-thermal-expansion issues that Cerebras had already solved in 2018; pioneering work’s reward is meeting the industry’s future problems first.

  • The cost was an 18-month stretch burning $8 million monthly without a solution, with board meetings every six weeks and the deficit climbing by successive hundreds of millions. Feldman’s operating lesson was that a startup is “a pressure test on your soul,” and survival requires modulating highs and lows while remaining a “professional David in the battle with Goliath.”

  • That experience informed his rejection of the idea that Elon Musk or Mark Zuckerberg can simply purchase AI leadership. Intel and AMD destroyed capital trying to enter mobile despite talent and fabs; skill is necessary but insufficient, while hard work, grit, ethics, organizational purpose, and culture make good luck more likely.

18. Compute gets cheaper, but power, trust, and orbit remain bottlenecks

  • Feldman resisted predicting Cerebras’ ultimate application despite a roughly four-trillion-transistor WSE3. Infrastructure companies “build roads”: his earlier Ethernet work helped drive network costs low enough for others to invent WhatsApp, and Cerebras similarly bets that sparse linear algebra and faster calculations will let frontier builders create uses it cannot foresee.

  • Orbital data centers fit his seven-to-ten-year category because “the last 10% don’t take 10%, they take 90%.” Cerebras has advantages—less chip-to-chip communication and the ability to disable damaged cores and route around radiation faults—but launch, software orchestration, shielding, and cluster communications leave production “the better part of a decade” away.

  • Ismail still expects a “10-cent lawyer” because learning curves, competition, capital chasing bottlenecks, and AI recursively optimizing chips and models drive intelligence costs down. He argued that routine legal and accounting work is vulnerable because it mediates obscure knowledge. Feldman disagreed with reducing lawyers to gatekeepers, emphasizing counsel and judgment, while acknowledging that routine document drafting can be automated.

  • The discussion cast China’s stronger hand as power infrastructure, not leading-edge compute: a panelist said China upgraded its grid while the US remained tied to a 1950s-era system. The episode also cited reports of Chinese proxy services offering American-model tokens at roughly 10-times discounts to capture reasoning traces for training—showing that token access, governance, and trust matter alongside nominal capacity.