Pioneers Insight Method Research Author
What You Missed in AI This Week (Google, Apple, ChatGPT)
Back to Episodes

What You Missed in AI This Week (Google, Apple, ChatGPT)

Summary

  • Google’s Veo 3 delivered what Olivia Moore calls “the ChatGPT moment for AI video,” pairing generated footage with native audio and multiple speaking characters in one prompt. That completeness helped drive million-view clips and “faceless channels” gaining hundreds of thousands of subscribers within days. The constraint remains severe: eight-second generations, no audio from image-to-video, and roughly $0.75 per generated second.
  • Voice models are competing on human imperfection, not merely intelligibility. ChatGPT’s upgraded Advanced Voice Mode now sounds more natural and expressive, with rising question inflections, filler sounds, and other human-like touches, after Sesame, Gemini, Grok, and NotebookLM made its original experience feel “not that advanced anymore.” ElevenLabs v3 pushes the same frontier into production tooling, using text tags to prompt whispers, emotions, sound effects, interruptions, and multiple characters.
  • Apple’s AI story remains cautious and disappointing to the hosts, with the most compelling announced feature being real-time translation across calls and FaceTime. Justine suggests Apple is outsourcing much of its “true AI” to ChatGPT while retrenching on an AI-native Siri after jumbled notification summaries caused backlash. The emblematic failure: Siri could not determine whether tomorrow was the month’s second Monday and instead offered to search ChatGPT.
  • Consumer AI has inverted the historical startup revenue curve: the median consumer company in a16z’s dataset reached $4.2 million of ARR after 12 months, versus $2.9 million at the bottom quartile and $8.7 million at the top quartile. Those figures were twice the corresponding AI-era B2B benchmarks, a sharp reversal from the pre-AI assumption that consumer companies would wait three to five years before monetizing. As Olivia Moore put it: “Consumers are back.”
  • Inference costs forced consumer startups to charge early, but product utility is supporting an average user payment of $22 per month—more than double the pre-AI subscription average cited by the hosts. Paid-user retention is roughly comparable with pre-AI consumer software despite heavy free-user “AI tourism.” Credit packs also introduce enterprise-like revenue expansion as power users spend another $10, $12, or $50 before their subscriptions renew.
  • Natural-language creative tools are collapsing brand development from a specialist workflow into a prompt-driven stack. Justine created the fictional Melt frozen-yogurt brand in under a couple of hours using ChatGPT for ideation, Ideogram for typography and packaging, and FLUX Kontext on Krea for consistent product and store imagery. Her larger call is that future entrepreneurs can assemble “full-stack AI brands”—including product design, apps, ads, avatars, influencers, and drop-shipped goods—without mastering tools such as Photoshop.

Deep dive

1. Veo 3 turns complete video scenes into one-prompt outputs

  • Olivia’s framing is categorical: Veo 3 was “the ChatGPT moment for AI video.” Veo 2 had established higher-quality scenes, consistent characters, and physics; Veo 3 added native audio, letting one text prompt generate dialogue and multiple speaking characters alongside the footage.

  • That completeness helped explain the distribution jump. Veo 3 clips attracted millions of views, while channels composed entirely of generated videos gained hundreds of thousands of subscribers within days. Stormtrooper vlogs work because masked characters, yetis, and capybaras are less sensitive to small visual discontinuities, while the model may already know what those characters or creatures look like.

  • The hard boundary is eight seconds. Audio works only from text-to-video, not image-to-video, making longer human stories difficult to maintain consistently; creators therefore use masks or nonhuman faces. The result is a new class of “faceless channels” that no longer requires a creator on camera.

  • Access also shifted quickly: launch required Google’s $250-per-month AI Ultra plan through Flow, while the model later became available via API and on platforms including Hedra and Krea through plans around $10 monthly, or through fal and Replicate on usage pricing. At roughly $0.75 per second, Olivia expects Google to pursue larger, longer-video models but says coherence and pricing will be challenges; she hopes to see more optimized or distilled models at lower cost.

2. Synthetic voice gets human through hesitation, emotion, and interruption

  • The upgraded ChatGPT Advanced Voice Mode sounds more natural and expressive, including rising question inflections and “um”/“uh” sounds. After NotebookLM, Sesame, Gemini, Grok, and open-source providers raised expectations, the hosts joked that it had gone “from Advanced Voice Mode to basic voice mode to Advanced Voice Mode”—and is finally advanced again.

  • Their unresolved question is why this took perhaps six-plus months while other providers advanced. One hypothesis is caution after controversy around “Her” and fears of a companion replacing human relationships; the other is prioritization across text-based AGI, Sora, images, and reasoning.

  • ElevenLabs v3 turns performance direction into editable text tags. Instead of recording someone crying, whispering, or using an accent and then converting that performance, creators can tag a line as “sadly,” “resigned,” or “whispering,” add sound effects, and script one character interrupting another.

  • Justine’s farm demo combines a thick Texas accent, mooing cows, and a second character cutting off Austin for allegedly faking the accent. Her key claim is that prompted interruption makes ads and narrative dialogue sound like “a natural conversation, which we’ve never had with AI voice before.”

3. Apple’s strongest AI feature is translation, not an intelligent Siri

  • The hosts remain disappointed by Apple Intelligence because the awaited “true personal assistant on mobile” has not arrived. Siri’s inability to answer what Monday of the month tomorrow would be—before offering ChatGPT search—captures the gap between basic user expectations and Apple’s current system.

  • Justine suggested that Apple seems to be outsourcing many substantive AI tasks to ChatGPT and retrenching after notification summaries grouped several notifications together and got jumbled. Apple highlighted Genmoji and call transcription; real-time translation across calls and FaceTime stood out as the “natural and obvious use case” with the clearest immediate value.

  • One host is still “holding out hope” for Genmoji after Olivia saw a viral Gen Z TikTok featuring Genmojis. The contrast is telling: Apple showed incremental interface features while the episode’s other models opened new forms of media production.

4. Consumer AI reaches millions in revenue faster than AI-era B2B benchmarks

  • The hosts’ analysis covered companies a16z met during roughly the first 22 to 24 months of generative AI, measuring growth from the start of monetization. The median consumer company reached $4.2 million in annualized revenue at month 12; the bottom quartile reached $2.9 million and the top quartile $8.7 million.

  • Those benchmarks are twice the comparable AI-era B2B numbers. Pre-AI, $1 million of first-year ARR was “amazing, best-in-class” for a B2B startup selling to enterprises, while consumer startups commonly spent three to five years building users before monetizing through advertising, marketplace transactions, or other later revenue.

  • The mechanism begins with inference costs: unlike conventional software, each additional AI query can cost cents or dollars, leaving an active user costing dozens of dollars monthly. Companies were “kind of forced” to charge, then discovered consumers would pay an average of $22 per month—more than double the cited pre-AI subscription average.

  • Utility supports that willingness across creative work, companionship, language learning, reading instruction, nutrition, and coaching. A $22 subscription can replace or broaden access to services that previously required a human charging $50 an hour, if that; vision models, for example, can turn meal photos into calorie and protein estimates plus daily or weekly dietary analysis.

  • Free users exhibit substantial “AI tourism,” but median paid-user retention is roughly as strong as in pre-AI consumer companies. Credit exhaustion creates upsells of $10, $12, or $50, while companies such as ElevenLabs can move from a $10 individual plan into high-ACV enterprise contracts—far faster than Canva’s cited five-to-seven-year consumer-to-enterprise journey.

5. Prompt-driven tools can assemble an entire brand in hours

  • FLUX Kontext’s distinction, in Justine’s telling, is preservation: it can move a person, product, or logo into a new setting while maintaining much greater consistency than the GPT-4o image model. That makes natural-language editing—“Photoshop but with natural language prompts”—usable for product photography and marketing collateral.

  • Her Melt workflow began with ChatGPT for the frozen-yogurt concept, name, colors, and logo direction; Ideogram produced typography and a branded cup; and, at Krea, FLUX Kontext let her explore placing the cup in restaurants or parks, changing packaging colors, and creating an ube variation. She also generated a store image and superimposed the logo onto it.

  • The next experiment would animate those assets with Veo 3 or Higgsfield and test whether the model understands how frozen yogurt melts and how it would land if the cup were tossed in the air—whether it would “kind of plop” as it would in real life. Justine made the prototype in less than a couple of hours; Olivia said the branding looked more exciting than many professional brands and suggested this could support agency campaign mock-ups.

  • Olivia says this points toward “full-stack AI brands.” Justine then imagines AI-designed logos, product photos and perhaps products; vibe-coded websites or mobile apps; drop-shipping to consumers; and social ads featuring AI avatars, promoted by AI influencers generated by Veo 3 that do not actually exist. The enabling shift, in Justine’s view, is that “you can just ask for what you want in a text prompt” and iterate until you love the result.