Pioneers Insight Method Research Author
Google's Gemini 3 Is Here: A Special Early Look
Back to Episodes

Google's Gemini 3 Is Here: A Special Early Look

Summary

  • Gemini 3 Pro’s clearest measured leap is on Humanity’s Last Exam, rising from Gemini 2.5 Pro’s 21.6% to 37.5%. Google reported similarly large gains across more than a dozen benchmarks, while Josh Woodward cited “cracking the 1,500 Elo on LMArena.” Roose relayed Google’s categorical pitch: anything done with ChatGPT, Claude, or an older Gemini should work better here.
  • Google is trying to move the interface from chatbot answers to software generated on demand. Gemini 3 can turn a request into a custom interface, as in an interactive Vincent van Gogh tutorial or mortgage calculator, while an agent under testing is intended to organize an inbox and propose replies. Casey Newton’s caveat: the agent was demonstrated through only “a few animated GIFs.”
  • Distribution may matter as much as model quality: Gemini 3 is available this week in the Gemini app and Search’s AI Mode, and to developers in various products, though not yet in Docs or Gmail integrations. Demis Hassabis said Google targets the “Pareto frontier of cost to performance” because products such as AI Overviews must serve billions. Kevin Roose’s investor-relevant inference is that Google can deploy frontier capabilities without “melting” its servers.
  • Gemini 3 does not shorten Hassabis’s five-to-10-year AGI forecast. He called progress “dead on track” but still expects perhaps “one or two more” breakthroughs, alongside better reasoning, memory, and world-model work such as Simmer and Genie. Scaling may face diminishing returns, he conceded, yet still offers an “extremely good return on that investment.”
  • Hassabis rejected a simple Google-is-winning narrative, arguing that rate of progress matters more than a temporary lead. He called Google DeepMind the company’s “engine room,” with its AI work reaching Search, Maps, YouTube, Android, Workspace, and Gmail as well as AI-first products. Google’s installed base gives Alphabet many paths to usage and near-term revenue.
  • Hassabis sees a selective AI bubble, not a sector-wide mirage. “Multi-$10 billion” seed rounds for teams with “basically nothing” look bubbly, but he sees perhaps half a dozen to a dozen greenfield businesses—robotics, gaming, drug discovery, and Waymo—with potential to become “massive multi-$100 billion businesses.” His portfolio claim: Alphabet should be positioned to win whether things continue or investment retrenches.
  • Greater tool-calling and function-calling ability improves coding and reasoning but also raises cyber risk. Hassabis called Gemini 3 Google’s “most thoroughly tested” model, citing internal work, external testers, and safety institutes. Josh framed the product as a tool and proposed measuring “how many tasks we helped you complete in your day,” rather than focusing on companionship.

Deep dive

1. Gemini 3’s benchmark jump must translate into felt utility

  • Roose’s headline number: Gemini 2.5 Pro scored about 21.6% on Humanity’s Last Exam, versus 37.5% for Gemini 3 Pro. Google supplied more than a dozen benchmarks showing the new model “beats the old one handily.”

  • Woodward highlighted “cracking the 1,500 Elo on LMArena,” but called benchmarks proxies. What ultimately matters is user satisfaction, and Google finds the two measures are “still moving in the same direction.”

  • Newton’s pushback — worth keeping: ordinary chat may already feel solved, leaving users unable to formulate a request that exposes a frontier model’s advantage. Woodward answered with concision and clearer presentation; Hassabis emphasized reliability, a more pleasant style, and a “step change” in vibe coding.

2. Generated interfaces turn answers into disposable software

  • Woodward said Gemini 3 can reason across many steps without losing its train of thought and separately highlighted new generative interfaces. Google’s examples were an interactive Vincent van Gogh tutorial and a mortgage calculator for a home costing more than $1 million.

  • Coding received “a lot of investment,” with Google Antigravity offered as a showcase. Hassabis said front-end and vibe-coding performance had crossed a “threshold of usefulness” strong enough to pull him back into games programming for projects over Christmas.

  • The Gemini agent under testing is intended to inspect an inbox, organize related messages, and propose replies. Newton was eager but explicit about the evidence: the briefing offered only “a few animated GIFs.”

  • Availability begins this week in the Gemini app, Search’s side-tab AI Mode, and developer products. Google gave no timetable for bringing Gemini 3 into the widely used Docs or Gmail integrations.

3. Google’s distribution machine is becoming the competitive thesis

  • Bringing Gemini 3 to Search suggested to Roose that Google can serve it economically at enormous scale. Hassabis said efficiency and distillation are necessities for “extreme use cases” such as AI Overviews, with other Gemini 3-era family models in development.

  • Asked whether Google had retaken the lead, Hassabis declined the framing: in a “ferociously” competitive market, “the only important thing is your rate of progress.” The work now is converting Google DeepMind’s research into products, with the lab serving as Google’s “engine room.”

  • That engine-room strategy spans Gemini, NotebookLM, Maps, YouTube, Android, Search, Workspace, and Gmail. Newton supplied the uncomfortable counterpoint: AI Overviews appear to lift Google usage and revenue—“not working out for the rest of the web,” but working for Google.

  • Google is also giving U.S. college students one free year of paid Gemini access. Its repeated “learn anything” positioning sounded to the hosts like a homework euphemism and a classic “first hit for free” adoption play.

4. AGI still requires breakthroughs beyond scaling

  • Hassabis’s five-to-10-year AGI estimate remains unchanged: Gemini 3 is “dead on track,” not an unexpected acceleration. Consistent general intelligence may still require “one or two more” breakthroughs plus improvements in reasoning, memory, world models, and physical intelligence.

  • His scaling answer rejected the binary choice between exponential progress and zero progress. Returns may diminish rather than double every era, he said, while remaining “well worth doing” and producing an “extremely good return” on continued investment.

  • Tool use illustrates the capability-risk coupling. Better function calling improves coding and reasoning, but also makes the model more capable in “riskier things too, like cyber,” requiring cautious misuse testing with safety institutes and external testers.

5. Alphabet is positioned for both an AI boom and retrenchment

  • Hassabis’s “strictly my own opinion” was that only parts of AI are bubbly. Seed rounds reaching “multi-$10 billion” for talented teams with “basically nothing” may be an early warning, but that does not negate the technology’s underlying value.

  • Longer term, he sees half a dozen to a dozen possible “massive multi-$100 billion businesses” spanning robotics, gaming, Isomorphic’s drug discovery, and Waymo. Nearer-term returns could come from reorganizing existing multibillion-user products, alongside cloud revenue and TPUs.

  • The portfolio logic is deliberately two-sided: Alphabet intends to exploit continued expansion, yet Hassabis believes it would also be “best placed” after a retrenchment. For an immediate consumer hook, Woodward chose selfie-based image editing; Roose’s verdict was simpler: “Nano Banana will save Thanksgiving dinner.”