Pioneers Insight Method Research Author
(Preview) Google Starts Dancing, The Winners and Losers of Gemini Week, OpenAI Has an Advertising Problem
Back to Episodes

(Preview) Google Starts Dancing, The Winners and Losers of Gemini Week, OpenAI Has an Advertising Problem

Summary

  • Gemini 3’s benchmark sweep is meaningful validation of Google’s comeback, not proof that every user will experience the best model. Ben Thompson’s base assumption is that the scores are “broadly correct” because Google has the resources, experience, and integrated stack one would expect eventually to win; he still says he is neither positioned nor qualified to certify “best in class.”
  • Anthropic emerges as the week’s other winner because coding was the one benchmark Gemini did not lead. Thompson thinks Anthropic’s edge, dating to what he thinks was Claude 3.5, has been surprisingly persistent; Andrew Sharp says it has “a moat for now.”
  • Benchmark leadership is hard to interpret when teams can optimize for or cheat the test. Llama 4’s benchmark-specific release is Thompson’s exhibit: disguising a mediocre model through benchmark gaming means “heads need to roll.” A friend’s characterization of Google’s bureaucratic tendency is that it “builds to the benchmark,” leaving real-world translation unresolved.
  • Google’s strategic weapon is not merely Gemini 3 quality but a TPU-based cost structure that could be sustainably lower than rivals’. Model-chip co-design, cheaper specialized silicon, and freedom from NVIDIA margins compound Google’s enormous cash flow; Thompson calls that combination “pretty devastating,” though some scalability benefits remain theoretical here.
  • ChatGPT’s installed base may blunt a technically superior Gemini just as Windows’ ecosystem blunted the better Mac. Thompson mentions 800 or 900 million ChatGPT users, then says not to quote him on whether that figure is weekly, monthly, or daily. Both stress that changing habits is harder than winning a fresh market: “Even if Gemini is better,” ChatGPT is already good enough and widely used.
  • Release-week hype is outrunning users’ ability to distinguish increasingly capable models. In Thompson’s own comparisons, ChatGPT beat Gemini on UniFi bandwidth controls and turkey-brining timing, including the crucial skin-drying step; anecdotal as those are, they illustrate the difficulty of judging model differences alongside Sharp’s iPhone analogy of steadier, less spectacular gains.

Deep dive

1. Gemini 3 finally makes Google dance on its own terms

  • Sharp opens with the 2023 taunts that frame Gemini 3: Altman worried about a “lethargic search monopoly,” while Nadella wanted the “800-pound gorilla” to dance—and wanted credit for making it dance. Pichai’s mild reply, warning against “playing to someone else’s dance music,” has aged into Google’s comeback narrative.

  • Thompson’s benchmark problem is perceptual: AI is already so good for so many users that differences become difficult to see. “You can see down easier than you can see up,” and models may soon leave virtually everyone looking upward without a reliable sense of relative intelligence.

  • Capability also remains “very spiky”: a model can excel in one domain, hallucinate in another, or visibly pull from Reddit. Thompson therefore finds AI most interesting where no right answer exists and invention is valuable, or where correctness is objectively verifiable.

2. Coding preserves Anthropic’s lead while exposing benchmark games

  • On the reported suite, Gemini led every benchmark except coding, where Anthropic remained first. Thompson thinks Anthropic discovered something around 3.5 that “blew everyone’s mind,” producing a surprisingly durable lead and making it “one of the big winners of this week.”

  • Coding is unusually informative because “the code has to run”: it compiles or it does not, throws errors or it does not. That verifiability supports both credible measurement and reinforcement-learning loops that test outputs and improve post-training.

  • Public benchmarks invite distortion. Thompson cites Llama 4’s benchmark-specific version and condemns any team “disguising their failure by cheating”: the underlying mediocre model is one problem, but “the failure” and “the cover-up” together mean “heads need to roll.”

  • Google presents a subtler concern. A friend once characterized its bureaucratic weakness as building to benchmarks—“you get what you measure”—so dominant Gemini scores could reflect real superiority or KPI optimization. Thompson’s answer is not dismissal but continued real-world use.

3. Ordinary tasks puncture the launch-week narrative

  • Thompson’s model usage has already moved with experience: Grok 4 felt like “a big leap forward,” but browser friction and what seemed like declining quality pushed him back to a ChatGPT that seemed to be improving. His standard is sustained utility, not release-week excitement.

  • On a UniFi networking question, Thompson says ChatGPT explained how quality-of-service policies could limit guest bandwidth, then warned that software handling might slow the whole network. Gemini omitted that option and instead proposed separate access points connected through a 100-megabit router or switch—“number one, bizarre answer.”

  • Thompson considers turkey preparation a domain where he would like to think he is genuinely expert. In his comparison, ChatGPT covered thawing, wet brining, and leaving the bird uncovered so its skin dries and crisps; Gemini started too late and skipped drying, implying soggy skin. “You wanna talk about a confidence shaker.”

  • Sharp’s pushback is that these releases increasingly resemble mature iPhone launches: once-stunning leaps have become steadier and “less sexy.” Thompson agrees but adds that Google benefits from a powerful comeback narrative as well as structural reasons to expect technical leadership.

4. TPU economics turn Gemini’s quality into a larger threat

  • Thompson expected Google eventually to build the best model: it has worked on the problem longest, commands immense resources, and owns a fully integrated stack. Google invented the transformer, but OpenAI’s more modular Azure/NVIDIA setup nevertheless took the lead; Gemini now provides important validation for Google’s integrated approach.

  • TPUs are less programmable than GPUs but also simpler and cheaper to manufacture. Google can co-design training architecture and silicon, making architectural choices with TPU capabilities in mind; Thompson contrasts monolithic models with mixture-of-experts systems and the resulting communication and scaling tradeoffs.

  • Newer TPU designs may support stronger scaling across separate systems or facilities, although Thompson does not claim Gemini 3 validates that specific capability. What the release does validate is that coordination between Google’s model and hardware teams “is paying off.”

  • The economic consequence is a potential “sustainable cost advantage”: Google avoids NVIDIA’s margins, uses cheaper chips, and retains most of the secret sauce even with Broadcom’s assistance. Add cash flow that Thompson says is “better than ever,” and Sharp’s summary lands: more cash than almost anyone, potentially at lower cost than everyone else.

5. ChatGPT’s installed base may outweigh Google’s integration

  • Thompson corrects the standard Windows-versus-Mac history: DOS and its software ecosystem preceded the Mac, while Windows inherited compatibility with that installed base. The Mac introduced the earlier graphical interface and may have been better, but Windows was effectively first at the economically decisive software layer.

  • That history also revises the simplistic claim that modular systems always defeat integrated ones. Consumer markets reward integration because the user is the buyer and “there is no ceiling to the quality of the user experience”; enterprise purchasing gives spreadsheets and organizational constraints more power, even after SaaS improved bottoms-up adoption.

  • AI is not starting fresh. Thompson mentions 800 or 900 million ChatGPT users but immediately says not to quote him on whether the figure is weekly, monthly, or daily. Switching those users is harder than winning an unformed market, and Gemini being incrementally better may not be enough to change behavior.

  • Sharp’s pushback—worth keeping—is that Google’s true victories may be preserving Search while growing Cloud and YouTube, not converting ChatGPT users. Thompson’s closing analogy agrees: consumer AI may resemble the PC era, where the integrated Mac “was better, but it didn’t matter because everyone was already using Windows.”