Pioneers Insight Method Research Author
Back to Pioneers
Logan Kilpatrick
Researchers 3 Curated Dialogues

Logan Kilpatrick

Google DeepMind · AI Researcher

Frontier Insights

Frontier Thesis: Compute economics and architectural integration redefine model dominance. Massive volume growth (500T tokens/month) and proprietary TPU scaling transform infrastructure into Google’s decisive moat, where raw intelligence per dollar and per second—via ultra-fast, low-cost Flash models—collapses margins for standalone developer tooling.

Strategic Moves: Google is commoditizing foundational inference while absorbing orchestration into its stack through Antigravity and Agent APIs, pairing massive context windows with sub-second execution to capture enterprise distribution.

Key Risks: Ecosystem lock-in via co-optimized frameworks threatens developer portability; sustained pricing power remains unproven against commoditization, while strict runtime constraints cap early agent autonomy.

Key Views & Dialogues

The Model Eats the Scaffolding: DeepMind’s Logan Kilpatrick & Tulsee Doshi on 3.5 Flash, Omni & More

  • 🗓️ Date2026-05-20 | 🎙️ Show:The Cognitive Revolution

Gemini 3.5 Flash leads with Tulsee citing roughly 280 tokens per second, about three times faster than other large models and significantly cheaper, prioritizing intelligence per dollar and second. Google is shifting its AI stack toward Antigravity’s shared agent harness, while preserving harness diversity; the payoff is a cross-product feedback loop, but co-optimization could still raise switching costs and pricing power.

View Dialogue Notes & Key Takeaways
  • Google is leading with Gemini 3.5 Flash because it is emphasizing cost-adjusted performance alongside frontier quality. Tulsee Doshi says slower responses hurt live Search and Gemini experiments “even if the model is hugely better,” while Flash is roughly three times faster than other large models and significantly cheaper. She cited “I think 280 tokens per second” on Artificial Analysis—a speed that can make an agent finish before she cancels it.

  • The missing Ultra brand does not mean Google stopped scaling frontier capability. Logan Kilpatrick calls model naming “marketing”: Pro has grown larger and more powerful, while Deep Think adds test-time scaling; internally, Pro influences Flash, which influences Flash-Lite through distillation, and recipes can also scale upward. Google’s portfolio choice reflects its need to serve multiple products and billions of users, where Flash and Flash-Lite economics matter alongside demand for high-end quality.

  • The strategic center of Google’s AI stack is shifting from standalone models to “model-harness-product-symbiosis.” Antigravity is becoming shared agent infrastructure for Gemini Spark, AI Studio, the Agents API, and eventually more Google products—the new cross-product through-line after Gemini itself. Logan’s shorthand for the direction is “agents, agents, agents,” with the harness handling tool loops and orchestration that individual teams previously had to build.

  • A co-trained harness could deepen switching costs, but Google says it is explicitly preserving “harness diversity.” Labenz’s investor-relevant challenge was that optimized model-harness stacks might create silos, stickiness, and pricing power. The guests want both outcomes: a seamless full-stack Gemini experience that generates better training and evaluation data, and models that still generalize to enterprise customers’ own orchestration systems.

  • The durable operational advantage may be infrastructure standardization and the feedback loop across Google’s product estate. Logan says the AI stack now needs rewriting “every 12 to 18 months,” making common infrastructure essential; Tulsee says products expose rough edges that flow back into model evals and data. “The model eats the scaffolding” captures the cycle: capabilities once implemented in surrounding code progressively move into the model or shared harness.

  • Recursive self-improvement is already part of Gemini development, but Google’s near-term framing remains human-directed collaboration. Gemini can submit code changes, launch evaluations, propose research improvements, and support parallel ablations; one safety-and-alignment lead reportedly ran such experiments from her phone in a hot tub and produced a report within an hour. Logan says the near-term picture of an autonomous “ML intern” launching enormously expensive pre-training runs does not seem realistic: the opportunity cost of a wrong direction keeps “the human in the driver’s seat.”

  • Google describes Gemini as a collaborator while treating apparent psychological distress as a model or product defect and questioning whether welfare interviews reveal an inner state. The team evaluates sycophancy, role-play, looping, and rabbit-holing checkpoint by checkpoint; Tulsee calls models going off the rails a “model bug.” She is skeptical of welfare interviews that ask models about deployments they cannot observe, arguing that without such context they are largely “pontificating.”

  • Rather than maximizing raw context length, Google is betting on search grounding, compaction, and speed. A 1 million-token request can cost “a few dollars,” limiting demand, while much of a giant context is distracting; the preferred direction is selecting the right information at the right time. Diffusion coding remains active research, but 3.5 Flash already delivers enough speed that Google is testing where further acceleration actually has diminishing returns.

  • 🔗 Original source & video: The Model Eats the Scaffolding: DeepMind’s Logan Kilpatrick & Tulsee Doshi on 3.5 Flash, Omni & More

Listen to full conversation →


The Decade of May 15-22, 2025: Google’s 50X AI Growth & Transformation with Logan Kilpatrick

  • 🗓️ Date2025-06-12 | 🎙️ Show:The Cognitive Revolution

Google’s AI usage rose 50× to roughly 500 trillion tokens monthly, strengthening the case that compute infrastructure, TPU scale, and organizational execution will increasingly separate frontier models. Application startups still benefit from speed and focus, while diffusion models and cheaper AI tools could create new software discontinuities even as frontier-model training consolidates.

View Dialogue Notes & Key Takeaways
  • Google’s AI usage rose 50× in a year—from roughly 10 trillion to 500 trillion tokens per month—amid an organizational rebuild, infrastructure scaling, and surging product demand. That is more than 50,000 monthly tokens per person on Earth, before counting other providers. Kilpatrick credits the mid-2023 Google Brain–DeepMind consolidation, repeated training-and-release cycles, TPU expansion, and DeepMind’s shift from foundational research into products; the key signal is “the slope of improvement” across all four.

  • Kilpatrick expects leading models to diverge as easy gains run out and structural advantages begin to dominate. Research initially converges because “people light the path,” while product teams race to avoid looking behind; from here, however, “the low-hanging fruit has been captured,” improvement gets harder, and Google’s compute infrastructure should matter more. Nathan’s pushback is the investor’s version: if DeepMind considers the next frontier difficult, tier-two foundation-model training looks brutal.

  • Application-layer economics remain unusually favorable even if frontier-model economics consolidate. Kilpatrick calls this “the best time in human history” to build a startup: billion-user incumbents solve general problems, while startups can attack narrow segments with speed, focus, and cheaper tooling. Nathan’s evidence is tangible—about $1,000 a month in AI subscriptions produces far more value, while one AI-assisted radio-ad project was priced roughly 75% below the incumbent workflow yet generated about $3,000 of project revenue per hour he personally spent.

  • The Windsurf episode exposes supplier risk, but Kilpatrick sees Google’s incentives favoring broad and early API distribution. He considers Anthropic’s decision to cut Windsurf off after its OpenAI partnership defensible as compute allocation to likely long-term partners. For Google, Cloud’s mandate, external evaluation signal, competitive momentum, and slow model migration inside billion-user products create “many levels of motivation and game theory” against withholding the best models—an expectation, not a Gemini 3 guarantee.

  • Gemini 2.5 Pro’s clearest differentiation is command of long context, not merely context-window size. Nathan’s 400,000–500,000-token codebase test felt like a step change; Kilpatrick cited OpenAI’s MRCR eight-needle evaluation, where the newest 2.5 Pro was “on the order of like 20% better.” His mechanism is that reasoning finally enables models to use the full window, pushing architectures toward more context loading even though RAG will remain necessary.

  • Reasoning models are turning agent architecture into a moving target: today’s necessary scaffolding can become tomorrow’s technical debt. Models are becoming “systems and agents out of the box,” while NotebookLM’s audio-overview pipeline has already fallen from 14 Gemini-heavy stages to four. The practical advice is to build structured workflows for present reliability while avoiding one-way-door designs that cannot be simplified as capabilities move into the model.

  • Diffusion language models could create another software discontinuity through sheer speed. Kilpatrick finds their coarse-to-fine generation more intuitive than strictly autoregressive writing and would not be surprised if the paradigm wins within two years. Nathan suspects diffusion could ultimately be better and notes that the transformer is not the end of history. Quality, cost, and performance trade-offs remain open, but Kilpatrick calls the demonstrated speed “unbelievably fast” and sees editing as a naturally strong fit.

  • Kilpatrick thinks AGI will be felt as a product assembly—and that authentic human perspective will remain scarce inside it. A model with perhaps 50% better reasoning, 50% better long context, and properly surfaced memory could produce the experience before any single release wins universal agreement as “AGI.” Yet he writes about 95% of his emails and posts without AI because agency and voice matter: “people want the Nathan experience,” even when NotebookLM can generate polished, topic-specific substitutes on demand.

  • 🔗 Original source & video: The Decade of May 15-22, 2025: Google’s 50X AI Growth & Transformation with Logan Kilpatrick

Listen to full conversation →


Breaking: Gemini 2.0 Flash Goes Live - Inside Google DeepMind’s Latest Release with Logan Kilpatrick

  • 🗓️ Date2025-02-06 | 🎙️ Show:The Cognitive Revolution

Google put Gemini 2.0 Flash into production at 10 cents per million input tokens and 40 cents per million output tokens, targeting text-to-app products with an estimated 40x reduction in LLM costs. Flash-Lite preserves the 7.5-cent price point while Pro pursues capability-first coding and reasoning, but Live API co-presence remains constrained by cost, memory, state, and a 10-minute session limit.

View Dialogue Notes & Key Takeaways
  • Google put the updated Gemini 2.0 Flash into production at 10 cents per million input tokens and 40 cents per million output tokens, while previewing Flash-Lite and experimental 2.0 Pro. Paid Flash defaults to 2,000 requests and 4 million tokens per minute, with higher tiers reaching 10,000 requests and 10 million tokens per minute. Logan Kilpatrick expects DeepMind’s tighter integration with products to accelerate both research and product progress by making the organization “end to end” from model creation through deployment.

  • Gemini’s strongest commercial wedge is cheap coding intelligence for text-to-app products such as Bolt.new and Cursor, with Lovable and v0 as possible adopters. Kilpatrick contrasted startups spending roughly $40,000-$50,000 monthly on LLMs with an estimated Flash bill of $1,000 or less—a “40x cost reduction”—and argued that falling costs particularly empower developers without tier-one VC backing. Google’s higher-end ambition remains categorical: “We’re going to have the world’s best coding model at Google,” with Pro and reasoning expected to carry that effort.

  • Flash-Lite primarily protects Google’s low-price positioning, while Pro is designed for capability-first experimentation. Flash-Lite preserves the 1.5 Flash price point of 7.5 cents per million tokens and does not support expensive features such as native image or audio generation; Nathan Labenz questioned whether moving from 7.5 to 10 cents could truly break a business. Kilpatrick largely agreed, conceding that continuity also prevents “Google’s raising the price for developers” from becoming the story. Pro’s price had not yet been announced because it was experimental, but Kilpatrick expected it to follow prior Pro models and be much more expensive.

  • The multimodal Live API points toward persistent AI “co-presence,” but today’s economics and architecture still impose a 10-minute session limit. Usage reports span coding help and blind users trying to navigate daily life. Separately, Labenz described using Advanced Voice Mode to guide his family through Mario 64. Kilpatrick expects iteration on cost, memory, state and context before always-present assistants can scale, but sees the direction as “very clear.”

  • Long context may become substantially more useful when paired with reasoning rather than immediate answer generation. A 1 million- or 2 million-token window can answer questions about a few relevant facts, Kilpatrick said, but synthesizing “a thousand different things” remains difficult; reasoning could let models work through the material and eventually move information via tools. Native image and audio output adds another vector: specialist image models may remain higher-quality in some domains, but Gemini can trade some quality for world knowledge—unlocking cases where prior models “weren’t smart; they were just good at generating images.”

  • Kilpatrick’s startup map centers on vision-language systems, reasoning-enabled agents, agent-facing internet infrastructure and better evals. He sees domain-specific computer-vision stacks as “up for grabs,” predicts reasoning will make more currently brittle agent products work over the next two years, and expects websites to need new protections, attribution and value-capture mechanisms once non-human visitors become normal. The evaluation problem remains difficult: benchmark sprawl obscures model choice, while even sophisticated teams struggle to turn taste into repeatable tests—“most of the problems in life end up being eval problems.”

  • 🔗 Original source & video: Breaking: Gemini 2.0 Flash Goes Live - Inside Google DeepMind’s Latest Release with Logan Kilpatrick

Listen to full conversation →