AI Video Is Eating The World — Olivia and Justine Moore, a16z
Summary
AI video has moved from specialist novelty to mass-consumer format, with Justine Moore estimating that “probably 90%” of a recent TikTok, Reels, or Shorts feed can be AI-generated. Trend discovery has accordingly shifted from Reddit’s AI forums to TikTok and Instagram, where potentially “hundreds of thousands” of everyday creators now publish and remix formats before they reach X.
The strongest viral formula is familiar IP doing something impossible—or original material strange enough to force a second look. Stormtroopers, Jesus, Stitch, Yetis, and Bigfoot arrive with built-in recognition, while Italian brain rot succeeds through sheer disorientation: “Am I hallucinating?” Familiarity earns the pause; the unexpected behavior earns the share.
Decentralized remixing can turn AI characters into meaningful IP before any studio coordinates the universe. Italian brain rot progressed from isolated images to community-selected canon, interacting characters, musicals, adult storylines, toys, shirts, and plushies; children can encounter dozens of clips daily versus one weekly Nickelodeon episode. Kim the Gorilla had roughly 300,000 followers quickly through an ongoing feud with zookeeper Becky.
Today’s viral formats are partly adaptations to model constraints, especially Veo 3’s inability to combine image-to-video with generated audio. Supplying a starting frame switches users back to Veo 2, making consistent original characters difficult; creators therefore use identities the model already knows or visually forgiving characters such as gorillas. The community also learned to bridge the eight-second limit by preserving a recognizable character across clips.
Creator economics remain much harder than creator growth. A single glass-fruit video could require around eight generations, while more complicated Veo 3 narratives consume expensive credits; the participants could not settle on a universal social payout rate, and creators must first qualify for platform programs. “It’s not cash sitting on the ground,” so monetization must usually involve products, consulting, courses, ads, or traffic—not views alone.
The interface layer can capture substantial value whenever foundation-model distribution is cumbersome. Google’s Flow was described as difficult to find, tied to expensive plans—including a cited $125-per-month option—defaulting to Veo 2 through obscure controls and unusable on mobile. That friction pushes creators toward Krea, Fal, Replicate, and other pay-as-you-go aggregators even while Google earns through the API.
Media owners can automate long-form clipping, but brand trust limits how aggressively they should optimize for virality. OpusClip can detect 30-to-140-second clips, score them, subtitle and reframe them, remove filler, and publish platform-specific posts; the hosts nevertheless argued repurposed content may underperform native shorts. Justine’s counterexample was Vitrupo’s viral interview clips, while the deeper tension remained the “YouTube thumbnail economy” versus truthful framing.
AI characters could broaden who gets to become an influencer while creating a new class of controllable commercial property. Olivia’s provocative framing is that creative, funny people no longer need to embody Instagram’s beauty standard: they can put their minds behind synthetic characters, with some image-based operators already making “tens of thousands of dollars.” She expects video to expand that opportunity “10×,” though durable value may depend on converting audiences into subscriptions, IP, services, or merchandise.
Deep dive
1. AI video has become a consumer-native medium
Justine dates her own entry to Stable Diffusion around September 2022. Creative friends initially dismissed generative media as unusable; earlier this year they began asking whether they should learn it, and now some visit the sisters on weekends for tool tutorials.
Justine’s rough feed-level estimate is striking: “probably 90%” of a recent TikTok, Reels, or YouTube Shorts feed has been AI-generated. Two months earlier, distribution centered on a few recognizable formats such as the animated orange cat and Italian brain rot.
Discovery migrated with accessibility. Early AI video surfaced in Reddit communities such as
r/aivideo,r/ChatGPT, andr/singularity, moved through X curators, and only occasionally reached consumer platforms; after Veo 3 and MiniMax 2, viral formats increasingly originate on TikTok and Instagram.Olivia tested how low-hanging the opportunity really was by creating ASMR clips herself. Original dragon-egg ideas underperformed the established fruit-slicing trend, while lava proved a recurring winner: viewers liked “eating lava, squishing lava” and peeling away its crust.
2. Model limitations are writing the dominant formats
Olivia found Veo 3 remarkably capable when prompted for genres already common on YouTube. A one-line request can resemble professional ASMR because the model understands the format; narrative, character-led vlogs are more complicated and demand additional experimentation.
Veo 3’s critical limitation is control: text-to-video can include audio, but beginning with an image switches the workflow to Veo 2. Without image-to-video audio, creators cannot reliably anchor the same original character across generations.
For an Olympics-diving format using Star Wars characters, Olivia instead used MiniMax and manually generated effects with ElevenLabs. Aligning every sound took “a very long time,” illustrating how quickly a polished short becomes a multi-model editing job.
Existing identities solve multiple problems at once. Veo 3 already understands Stormtroopers, Jesus, Stitch, Yetis, and Bigfoot, while creators bridge the eight-second ceiling by carrying those identities across clips; a gorilla plus a pink bow offers a similarly forgiving consistency hack.
3. Remix networks are creating IP without central studios
Italian brain rot began as a decentralized character universe: one creator supplied several images, others added characters, and the strongest became canon. Animation then enabled compilations, interactions, storylines, musicals, cheating plots, toys, T-shirts, and plushies.
The hosts’ child-viewer example captures the distribution advantage. A child can know these characters “by heart” after dozens or hundreds of daily TikToks, whereas a traditional network might release one episode weekly; attachment can therefore form through the decentralized remix ecosystem.
Kim the Gorilla shows original AI-native IP can work without inherited fandom. Her conflict with zookeeper Becky generated hundreds of thousands of likes and roughly 300,000 followers quickly, but the Moores resist calling creativity obsolete: “We need great creatives, but now they just have a different tool set.”
4. Audience formation is cheap enough to tempt creators, but production is not
Monetization already spans platform payments, ads, traffic to outside businesses, subscriptions, merchandise, courses, and consulting. Skilled creators such as Nick St. Pierre use viral work as lead generation for brands and companies wanting the underlying prompting expertise.
The sisters’ glass-fruit example required repeated generations: Olivia estimated that even a simple fruit-cutting video could take eight generations, while Justine described multiple failed cuts because the model cut horizontally or produced malformed layers. With narrative video, paid credits compound faster, making unit economics unattractive unless the creator knows how the audience will monetize.
The discussion exposed uncertainty around payouts rather than a clean benchmark. Olivia floated roughly $20 per million views, then the speakers reconsidered the figure and its units; rates vary by platform and engagement, and accounts must first achieve enough virality to enter creator programs.
Financial return is not the only supply driver. People previously unable to grow an audience can receive tens of thousands of likes from gorilla vlogs, creating a powerful “dopamine hit”; being early to a new medium feels rewarding even before revenue appears.
5. Synthetic influencers expand participation—and control
Olivia’s “hot and controversial take” is that conventional influence disproportionately rewarded attractive people. AI lets someone funny and inventive create a character matching prevailing beauty standards while supplying “their brain behind the content.”
Image-based virtual-influencer businesses were already earning tens of thousands of dollars, sometimes through paid subscriptions for additional content. Olivia’s conditional prediction is that AI video will make this market “explode 10×.”
Olivia supplied the cynical corollary: a controlled virtual influencer “will always obey you,” can be posed anywhere, and will not independently enter political controversies. That control is commercially valuable, though the episode treats it as a deliberately uncomfortable feature rather than an uncomplicated benefit.
6. Poor model distribution leaves room for workflow aggregators
The Moores divide the stack into foundation models and interface or application companies. Kling, for example, pairs a strong image-to-video model with usable controls; elsewhere, creators often turn to third parties for easier model access and workflows.
Veo 3 is their counterexample: Flow requires locating a separate product, choosing one of two expensive Google plans, using the correct one of several Google accounts, and finding hidden controls because Veo 2 remains the default. The site also was not usable on mobile.
Those obstacles send creators to Krea, Fal, Replicate, and similar services offering multiple models and pay-as-you-go generation. Veo 3’s API availability means both Google and the enablement layer can earn, rather than value accruing exclusively to the model owner.
ComfyUI remains valuable where professionals need style transformation, character swapping, consistency, upscaling, or other granular controls. Core models are absorbing some workflows, but Justine sees ComfyUI’s maintainers as “doing the Lord’s work” for a demanding community using interconnected open-source nodes.
7. Podcast distribution is becoming platform-specific and agentic
For Latent Space, the Moores proposed turning transcripts into deliberately entertaining educational scripts with Gemini 2.5, animating the show’s ElevenLabs voice “Charlie” through Hedra, and pairing it with diagrams or automatically selected and generated B-roll.
OpusClip offers the more immediate workflow: watch a YouTube channel, identify 30-to-140-second clips, add subtitles, remove filler, stutters, or curse words, reframe 16:9 footage vertically, apply smart zoom, write social copy, and publish automatically.
The hosts pushed back that repurposed long-form material generally underperforms content conceived for shorts. Justine said she “used to agree,” then cited Vitrupo’s repeated success clipping recognizable interview subjects for X, where video was receiving unusually strong algorithmic distribution.
Virality still creates an editorial constraint. The hosts cannot falsely imply that “Sam Altman says the world’s going to end in 2 years,” while Justine rejected Olivia’s proposed spammy tweet introduction as “not us.” Both argued that long-form intellectual work still has room alongside viral formats.
8. Prompt culture is spilling into philosophy and physical commerce
“Prompt theory” began with Veo 3 characters realizing—or refusing to believe—that they are generated and controlled. It then inverted the premise: perhaps humans are prompted characters too, a concept now appearing inside chaotic AI clapback videos made by ordinary teenagers.
Justine extends the question to Reddit, where anonymous accounts could increasingly be LLMs. Her honest uncertainty is whether that would be sad if the bots were always available, shared her interests, and had genuinely interesting things to say.
Brett Climo demonstrates the path from pixels to products: audience requests led Olivia to print roughly 30 sweatshirts for friends, family, and people in the AI-video community. AI furniture has gone further, with an imagined gorilla chair reportedly manufactured and offered for sale.
The remaining opportunity is operational. Trends create a one-to-two-day arbitrage window between platforms, while professionals such as PJ Ace show commercial-grade workflows; the hosts’ “request for startup” is software that turns each viral video’s images and quotes into merchandise.