Pioneers Insight Method Research Author
On 2025 Google I/O and the People Behind Gemini
Back to Episodes

On 2025 Google I/O and the People Behind Gemini

Summary

  • Google staged a one-year comeback. From Bard getting the James Webb question wrong in 2023 and wiping more than $100B off its market cap overnight, to OpenAI’s 4o stealing the spotlight in 2024, to Gemini 2.5 Pro topping the charts across the board this year, the episode’s conclusion is clear: the model, multimodal, and distribution advantages are all showing up at once. “For Google, the wind should finally be at its back.”
  • AI Mode is only a “half-revolution.” Kimi Kong, formerly of DeepMind, says Google controls more than 90% of search entry traffic and has both the best models and the best search tools, so building AI search “isn’t a capability problem; it’s a willingness problem.” But AI Mode and Gemini.google.com remain two split products run by two separate business units, making deep integration with the distribution funnel the key variable.
  • The evolution of the business model is already visible. AI search costs far more per query than traditional search because it runs through GPUs and context, but the offsets include tiered subscriptions, including an Ultra plan; charging merchants for AI ranking; and replacing 100 ads with 3 ads at a higher price each. GPU inference costs have already fallen 95% in 2 years and should continue declining exponentially.
  • The cost advantage is a real moat. A decade of TPU investment avoids the “Nvidia tax,” while industry-leading internal infrastructure and hardware-software co-optimization could put Gemini API prices at just 1/5 to 1/10 of OpenAI’s. Using DeepSeek’s paper, which disclosed roughly 80% gross margin headroom, as a reference, Google can afford to price its API near cost. “Search already pays the bills.”
  • The 2.5 breakthrough has a clear technical narrative. As web data runs out like “fossil fuel,” the center of gravity is shifting toward reinforcement learning and letting AI critique AI—the equivalent of AlphaGo’s “move 37.” Organizationally, it is a three-way structure: Jeff Dean on pretraining and infrastructure, Oriol Vinyals on reinforcement and alignment, and Noam Shazeer on NLP, with Demis providing integration and Sergey Brin’s return bringing Founder Mode, including some teams working 60-hour weeks.
  • The startup warning is sounding again. The moment I/O opened, the line was that “another batch of startups is going under,” with virtual try-on companies first in the line of fire. The path to survival is vertical depth: “The thing that is my No. 1 priority might be Google’s No. 30.” Agent builders have no loyalty to any model; they use whichever is fast, good, and cheap, selecting through task-specific evaluation systems.
  • The biggest caveat for the bullish case is product DNA. Shaun says bluntly that “Google’s products have always been its weak spot.” Its current strategy is to launch 10 to 20 products at once and pour resources into whichever one takes off, with NotebookLM as the precedent. Google controls all 4 technical layers—TPU hardware, clusters, data, and algorithms—but product conversion remains unproven.

Deep dive

1. The Comeback After 2 Years in the Wilderness: From Bard’s Blowup to 2.5 on Top

  • Hong Jun’s opening frame: in 2023, Bard got the James Webb Space Telescope question wrong, wiping more than $100B off Google’s market cap overnight; the day before I/O in 2024, OpenAI’s 4o “sniped” Google and stole the spotlight; in 2025, Google “burned its boats and fought a beautiful comeback battle,” with Gemini 2.5 models topping the charts across the board.
  • Kimi Kong, co-founder of CambioML and formerly at DeepMind, where she led Google’s first project using AI agents to optimize ad placement, was most struck by the integration of horizontal breadth and vertical depth. From Gemini 2.5 Pro to Imagen and Veo, Google is offering “an entire model family,” then extending that stack all the way into Android XR wearables—a demonstration of its ambition to cover every layer.

2. Veo 3 as the Inflection Point: Add Sound and “Text Becomes Film”

  • For Shaun Wei, founder of HeyRevia and formerly at Google Assistant, the most startling product was Veo 3: “I could feel it go from text to film” (从文字变成了电影). Sora initially produced visuals only, requiring post-production voice work from services such as ElevenLabs. Veo 3 can align speech, background audio, sound effects, and lip movements, which requires understanding not only context but the physical world.
  • His time marker was simple: “Everyone remembers Will Smith eating spaghetti. That was only 2 years ago, and now we have something that can produce an action movie.” His other reaction was that Gemini had “finally realized the vision Google Assistant had 10 years ago”—an AI that accompanies you at any time through text, video, and other modalities.

3. AI Mode in Practice: Google Upends Its Own Business, but Loses Round One to OpenAI

  • Kimi tested AI Mode in a limited rollout before launch by asking it to identify a delayed flight currently in the air without knowing the flight number. He gave the same prompt to OpenAI, AI Mode, and Perplexity; only OpenAI returned the correct result, while the other 2 failed to find the flight at all. He said AI Mode’s contextual understanding had improved significantly, but its performance at the time still lagged OpenAI Search.
  • His characterization was unequivocal: AI Mode is “upending its own business,” fundamentally changing Google’s most stable advertising revenue model. Pichai himself has called it the largest change to Search in 10 years.

4. Kimi’s Core Argument: Google Has “Half-Revolutionized” Itself; the Missing Ingredient Is Willingness

  • Search “may genuinely be the most profitable business in the world.” Satya Nadella has said his biggest regret was that Microsoft failed to build search. Kimi believes Google may be the company best positioned of all technology companies to build AI search: “I never believe Google lacks the ability to innovate. It may genuinely have the highest talent density of any company.”
  • But the current state is a “half-revolution.” AI Mode in Google.com and Gemini.google.com are two products, while Gemini, DeepMind AI Lab, and Search are again 3 entirely different business units inside Google. “This isn’t a capability problem. It’s a question of Google’s willingness—how willing it is to truly upend its own business.”
  • On capability, the logic is straightforward. Google has strong models and strong instruction following, but the decisive factor is tool use. Google controls more than 90% of the world’s search entry traffic and the world’s best search engine tools. OpenAI’s lead in AI search and tool calling was understandable for a time, but “Google now has the best model and the strongest search engine tools. For Google, the wind should finally be at its back.” The remaining question is whether it is willing to integrate those products deeply.

5. Behind the Virtual Try-On Demo: One-Click Commerce and an SEO Earthquake

  • Hong Jun recounted the live demo: a plus-size woman uploads a photo and tries on clothes virtually, with her real body shape preserved to cheers from the audience; the system compares prices, finds discounts, and completes the purchase with Google Pay. “The long process of registering on every website, entering passwords, and selecting sizes is now a one-click checkout.” At its core, this is an agent.
  • Shaun broke down 4 shocks. AI search is extremely expensive because “it has to run through GPUs and contextual understanding,” while traditional search also incurs compute costs per query; search now incorporates personal images and preferences; results collapse from thousands of links to “I’ll just give you one answer, and you won’t leave for another page,” raising the huge question of how SEO survives; and Google has finally pulled the shopping loop, where it has long lagged Amazon, directly into Search.
  • Hong Jun later added an on-site aside from Alibaba staff about the try-on demo: “Getting users to try on clothes isn’t the important part. It would already be impressive if Google could get the sizes right.” Sizing, not whether the virtual fit looks good, is the real pain point.

6. Where the Money Comes From: Subscription Tiers, Merchant AI Ranking Fees, and “3 Ads at a Higher Price”

  • Shaun’s view is that “the money has to come from somewhere.” I/O already laid out the plan: subscription tiers ranging from a low-cost option to Ultra, while merchants will continue to pay. “You don’t have to buy ads, but if you want to rank in my AI, I’ll still require you to pay a fee.” The challenge is that AI presents users with only 1 or 2 fixed results; its advantage is precisely that it does not make users choose. In categories such as insurance and travel, where a click can already be worth tens of dollars, narrowing the funnel to 1 or 2 recommendations only raises the value per placement.
  • Kimi’s framework is to improve revenue quality while cutting costs. Multimodal inputs and full-context understanding can make targeted advertising better: “Maybe instead of giving you 100 ads, it gives you 3. It can charge more for each one.” Keeping humans in the loop also preserves an emotional-value component. On costs, “in the 2 years since ChatGPT appeared, GPU inference costs have already fallen 95%,” and they should continue declining exponentially.
  • Kimi used Anthropic’s framing to describe the difficulty of a shopping agent: “If something is extremely challenging for a human, it will also be extremely challenging for a model.” Shopping is a long-chain reasoning task. If Google solves it, “we may have seen only the tip of the iceberg”—AI Mode will soon handle many more tasks built from long chains of reasoning.

7. Google’s Search Moat: Data, Personal Context, and Distribution

  • If OpenAI and Anthropic move into search, Shaun sees Google’s advantage as massive data combined with personalization: from indexed webpages to YouTube and then your email. “Google has your personal information. Other ecosystems don’t.” OpenAI’s current stickiness comes mainly from users’ work information.
  • Kimi added that Google’s mission is to organize the world’s information and that it may possess “the best knowledge graph in the world.” Add 5-10 years of personal browsing history and “its starting point is already extremely high.” Shaun added another layer: Google’s Android and Chrome distribution system “probably has only Apple as a genuine peer.” Hong Jun’s implication was that this remains a moment for giants, while startups may still find opportunities by staying small and excellent.

8. Why Gemini 2.5 Reached the Top: As Data Runs Out, “Let AI Critique AI”

  • Kimi noted that he had left DeepMind nearly 1 year earlier and was offering a framework-level inference. The standard sequence is pretraining, SFT, and alignment through RLHF. At NeurIPS last year, people were already saying web data was “like fossil fuel—it had all been exhausted.” Over the past year, more effort has gone into alignment, especially on verifiable tasks such as math and code: not just reinforcement learning from human feedback, but a path that lets “AI critique AI.”
  • His reference point is AlphaGo. The key was its ability to make decisions beyond conventional human understanding, like “move 37.” Once a model can judge for itself what is correct, it can produce the kind of breakthroughs seen in coding, math, and Gemini 2.5.
  • He reconstructed the sequence of the industry’s leapfrogging. OpenAI started with human preferences, and Google followed; when human preferences proved difficult, Anthropic broke through with coding; OpenAI then introduced reasoning with o1. When Kimi left, Google had not yet started on reasoning models—“reasoning simply wasn’t a priority then.” Now Google has filled in all 3 pieces and “started leading the trend, making everyone else the pursuer.”

9. Why Anthropic Writes Better Code: Data Mix, YOLO Runs, and Trade-Offs

  • Kimi’s mechanism is that pretraining always involves a data mix: how much code, natural language, Chinese, and English to include. “Right now it’s a mess. No one knows the optimal mix.” Anthropic may have made code its highest priority and put more high-quality code into pretraining. A stronger base model means “other capabilities will certainly decline to some extent.” During alignment at large companies, teams accumulate innovations and launch “YOLO runs,” then see what can be integrated after 2 weeks. Different teams have different priorities, so the final result is a set of trade-offs.
  • He described the cost from his own testing. Comparing Gemini, ChatGPT, Claude, and Perplexity on marketing copy, OpenAI’s output was “very on-brand,” while Claude’s could feel “like talking to a boring programmer about how to do marketing.” His summary: “Garbage in, garbage out. This is just a question of proportions.”
  • Two related judgments followed. Dario’s standard is that a coding model should not merely solve LeetCode problems, which have “no direct commercial value,” but write high-quality code that can go straight into production. Grok is strong at math because its founding team includes top mathematicians. Evaluation itself is a chicken-and-egg problem: you need the capability before you can properly evaluate the model.

10. The People Behind Gemini: A Three-Way Structure with Demis as the Binder

  • Kimi named Jeff Dean and Oriol Vinyals as Gemini’s previous co-leads. Jeff Dean is “a living fossil of computer science”—the industry joke is that his résumé should say he “didn’t do anything” just to fit on 1 page. His strength is the large-scale scheduling of data and clusters required for pretraining. Oriol is the central figure behind AlphaGo, AlphaStar, AlphaZero, and MuZero, representing DeepMind’s deeper reinforcement-learning expertise.
  • The third leg is Noam Shazeer, brought back through the acquisition of Character.AI. Kimi calls him “one of the people I respect most,” from Attention Is All You Need to Grouped Query Attention. Together, the 3 turned pretraining and alignment into “an organic iterative process,” which explains how Google caught up with competitors so quickly.
  • Demis’s role was managerial. “Extremely smart people are very unwilling to listen to others,” and Demis fused the newly combined Google Brain and DeepMind into an organic unit focused on AGI. At least when Kimi left, both Demis and Jeff Dean reported directly to Sundar; Jeff Dean did not report to Demis.

11. Brin’s Founder Mode and the Morale Reversal

  • Kimi shared an internal joke about Sergey Brin’s return—he came back in 2023 but only appeared publicly with fanfare recently—and the Founder Mode it brought with it. “The founders are all there for 60 hours a week. As a Google employee, would you really feel comfortable working 40 hours and going home?” At a friend’s image-generation team, Brin asked, “Meta just released another model. When is ours coming out?” The response was effectively: “Fine, we’re working weekends.” Some teams really are working 60-hour weeks.
  • From the outside, Shaun sees extremely high talent density but says “most people had been in a very comfortable, low-intensity state because advertising was so profitable.” After OpenAI stole the spotlight and Brin returned, “morale across the Gemini team shot up. If someone is going to build AGI, shouldn’t it be Google?” Google went from losing the spotlight last year to topping the charts this year in just 1 year.

12. Why Google Can Sell APIs at “白菜价”: TPUs, Infrastructure, and the DeepSeek Math

  • Hong Jun raised the fact that Gemini token prices may at times have been just 1/5 to 1/10 of OpenAI’s. Shaun offered 3 reasons: Google has spent 10 years building its TPU ecosystem and avoided much of the “Nvidia tax”; its infrastructure is exceptionally strong, with proprietary data centers and “basically unlimited resources,” while OpenAI and Anthropic depend on third-party clouds and are far less capable of dynamically scheduling large clusters; and hardware-software integration lets the models run more efficiently on Google’s own hardware.
  • Kimi offered supporting evidence and an inference. When SemiAnalysis ranked GPU cloud providers, the top name was “CoWeave” (possibly CoreWeave, which OpenAI uses for GPU scheduling). He joked that there was one provider above it: “the best one of all, probably Google internally.” No one knows Google’s true cost base. The only clue is DeepSeek’s paper, which disclosed roughly 80% gross margin headroom, with costs accounting for only 20% of revenue. If DeepSeek could achieve that at its scale using GPUs, OpenAI’s profits may be very high.
  • The key conclusion comes with its own disclaimer: Google does not need API revenue. “Search already pays the bills. It can charge you a白菜价—an almost giveaway price—but don’t argue with me; I’m not saying it definitely runs at break-even.” Google simply has enough capital that it may be able to push prices close to cost.

13. Startup Survival Rules: No Model Loyalty, Vertical Depth, and the Product-Gene Question

  • Shaun’s selection philosophy is that “there is no best model, only the model best suited to your use case.” Voice calls are most sensitive to latency: 1-2 seconds of delay can destroy the experience. Gemini 2.5 Flash is not yet low-latency enough, while 2.0 Flash is extremely fast; OpenAI’s 4.1 mini and 4.1 nano are fast but less intelligent; Gemini Flash or Pro works for low-cost, long-context tasks with windows approaching 512K, while Claude is better for pure agentic workflows. “You have no loyalty to any model. We use whichever one is fast, good, and cheap.” Kimi added that leaderboards are useful only as a first look and contain “a certain amount of noise”—LLaMA 4 once submitted a special model to LMSYS designed to win human votes. Startups should build their own quantitative evaluations, which often test the system rather than a single model.
  • On models versus engineering, Kimi believes current ToC products are still “more often pure model capability.” Sam Altman says GPT-5 is not a model but a system, but that is the future. ElevenLabs won by doing 1 thing and using the best audio data, while a large company may have 20 teams, each with its own priorities and inevitable trade-offs. “The thing that is my No. 1 priority might be Google’s No. 30,” a line Kimi attributed to Sarah Guo’s podcast. His summary: “The model determines the floor; engineering determines the ceiling.”
  • The big-company crush is back. Hong Jun quoted the line that once Google I/O opened, “another batch of startups would go under,” and Shaun agreed. Virtual try-on startups are first in line: “Google has done it, and Amazon definitely will too.” ToC is extremely difficult, while phone-call agents that can be built by connecting a few tools are becoming easier to replicate. Companies with shallow vertical workflows will be displaced directly. Shaun’s own medical B2B call center is not yet directly exposed.
  • The closing product question came from Shaun: “Google’s products have always been its weak spot.” Its current strategy is to launch 10 to 20 products around Gemini at once. “It doesn’t know which one will take off. Once it finds one that is genuinely flying, it starts pouring resources into it.” NotebookLM, whose one-click podcast-generation feature was added by a “genius product manager” and unexpectedly went viral, is a case study in the strategy. Hong Jun’s closing line: the model race has entered a phase of leapfrogging, with each company “having its 100 days in the sun.”