Victor Riparbelli, CEO @Synthesia: OpenAI vs Anthropic vs X.ai - Who Wins and Why | E1246
Summary
- AI is in a bubble whose real signal will be renewal, not bookings. Buyers told to produce an AI strategy will readily fund $50,000 pilots or spend $100,000 without knowing what they need, letting startups mistake budget availability for product-market fit. Riparbelli’s warning: “The real signal is renewal,” and weak products will meet a wall of churn as 12-month contracts expire.
- The application-layer defense is workflow breadth, not a prettier foundation model. Customers sign Synthesia’s million-dollar-plus contracts to move from an idea through drafting, editing, multilingual playback, publishing and distribution—not merely to generate an avatar. Calling OpenAI’s application-layer expansion an existential threat is “classical VC brain”: model capability matters, but durable value sits in solving the customer’s complete job.
- Foundation-model gains may continue, but scaling alone will not determine the winner. Riparbelli expects compute, algorithms and data all to matter; a 10–100x algorithmic efficiency breakthrough could overturn any linear “most compute wins” forecast. Given $10 million at the stated valuations—OpenAI at $160 billion, Anthropic at $60 billion or xAI at $50 billion—he would choose xAI for its asymmetric upside, X distribution and real-time data: “I would never bet against Elon.”
- Synthesia’s early capital scarcity helped force customer focus before the market understood generative video. After raising $1 million at a $5 million post-money valuation in 2017, the company failed over nine months to raise a planned $8 million and settled for $3.1 million. Riparbelli believes a larger round would have funded distractions such as deepfake detection; instead, the team charged from day one and learned that “you cannot use money to buy your way” to product-market fit.
- Riparbelli’s $50–100 billion Synthesia thesis treats text and slides—not conventional video production—as the addressable market. Once video and audio become as scalable as text, training, support, software sales and other communication can become visual and interactive. Capture even 5% of the world’s text communication, he argues, and “all the world’s communication is the market.”
- Generative content drives technical production costs toward zero while making judgment scarcer. Riparbelli expects more AI-made than human-made content within five years and eventually a Hollywood film’s purely technical production cost to approach zero, collapsing the production-value gap between creators and studios. Stebbings’ pushback is crucial: today’s clipping tools remain far below expert teams, and six different moments from one episode could each be “best”; Riparbelli concedes that storytelling, taste and asking the right questions become more important, not less.
- Trust shifts from proving content is “real” to proving who authorized it and how it changed. Riparbelli imagines Shazam-like fingerprinting that identifies an original work, its creator, date and subsequent edits, using a centralized database—though he said it could actually be a blockchain. Unverified material would “stick out like a sore thumb.” Synthesia chooses to police its own gray zone strictly because protecting enterprise customers’ brands matters more than collecting $30 monthly subscriptions from questionable conspiracy creators.
- AI adoption is gated more by people than raw model capability. Workers must trust, buy and integrate the technology, while governments can either preserve entrepreneurial incentives or damage the ecosystem; Riparbelli’s blunt warning is that “Europe has all the AI regulation and none of the AI companies.” He still gives the US West Coast a higher probability of startup success, but sees London’s loyal talent base and emerging global leaders as valuable if the UK avoids excessive taxation and regulation.
Deep dive
1. Capital scarcity helped Synthesia discover a business
Synthesia began in 2017 around an unfashionable premise: generative AI would move computing from analyzing existing data to creating video, speech, audio and music. Investors heard a claim that a Hollywood film would eventually require “nothing else than your imagination”; roughly 80–90 of them declined. Riparbelli partly attributed the resistance in Europe to a private-equity-oriented VC mindset that was less receptive to a technologist’s vision.
The eventual first round was $1 million at a $5 million post-money valuation, with Mark Cuban investing. Harry cited a $2.1 billion valuation and asked whether Cuban still owned roughly 15%; Riparbelli said it was “something in that range.” He later said he would choose Cuban again because he believed when almost nobody else did: “I’ll be forever grateful to him.”
A later attempt to raise $8 million became a nine-month process Riparbelli called “a big shit show.” With cash running out, Synthesia rewound the plan and raised $3.1 million, carrying it toward the Series A and an actual sustainable business.
Riparbelli thinks an easy $8 million would have widened the roadmap into deepfake detection, then the market’s dominant association with AI video. Constraints instead made the team “ferocious about charging people from day one,” even when the price was only £500, and kept attention on the commercial use case.
2. Product-market fit remains a recurring founder responsibility
Stebbings rejected product-market fit as a permanent state: winning creators says nothing about winning SMBs or enterprises. Riparbelli agreed that a successful company is “a long series of product-market fits,” although the first spark creates something employees can subsequently scale.
Before that spark, adding product managers, salespeople or 15 employees can slow learning rather than accelerate it. “You cannot use money to buy your way” to product-market fit; founders must personally understand the customer, and Synthesia needed roughly two years to understand video from first principles.
Mature product-market fits can become staffed businesses, but Riparbelli keeps the founder focused on the next market and where the company must be in two or three years. His recent relearning was to stop treating competitors and AI-market noise as the strongest signal: “Listen to the customer. They’re the ones who pay your bills.”
Synthesia had not touched its Series A before raising its B, its B before the C, or its C before the $100 million Series D led by NEA, with existing investors participating. Riparbelli still values a war chest for nine-month go-to-market ramp times and says capital supports the unit economics, revenue generation and $50–100 billion outcome he wants; bootstrapping all the way is a myth. Secondary liquidity reduces founder pressure: without it, he is “100% certain” Synthesia would have considered selling earlier.
3. Text replacement makes communication the addressable market
Riparbelli’s long-range thesis begins with text as a scalable but low-fidelity compression technology: translating thoughts into documents strips away context. Humans generally absorb richer information through voices, faces and visual signals, yet text dominates because it has been the only scalable medium for storing and sharing information.
Synthetic creation changes that constraint. Once video and audio no longer require cameras, microphones or physical capture, they can scale like text; Riparbelli even suggested that “our kids’ kids” might belong to one of the last generations that reads and writes by default.
TikTok provides the consumer preview, with video replacing almost every interface element and even video responses supplanting textual comments. Enterprise analogues could include interactive product demonstrations, training and support videos that answer spoken questions and immediately show the requested functionality.
Synthesia therefore defines its market as text and slides, not existing video production. “All the world’s communication is the market,” and converting even 5% of global text communication into video could, in Riparbelli’s view, support a company worth more than $100 billion.
4. Competition validated the category and accelerated its leader
For years, outsiders saw Synthesia as a “cute UK company” making amusing avatar clips, while management saw training and learning as merely the first wedge into a much larger market. That information gap kept the company’s growth and customer enthusiasm relatively hidden.
Competitors eventually arrived, including one that Riparbelli said copied Synthesia’s mission statement and market language “literally verbatim.” Annoying as that was, it supplied feedback unavailable to a category pioneer: management could observe rival experiments and learn from both their successes and mistakes.
Competitive marketing then educated buyers for the whole category. Riparbelli attributed part of Synthesia’s massive reacceleration in the second half of the year to rivals creating demand that flowed toward what he considered the superior product.
The new $100 million round inevitably signals strength to investors and rivals, but Riparbelli denied using capital primarily as intimidation. “I have never made a decision at Synthesia based on what our competitors do or don’t do”; capital only helps when management already knows where to deploy it.
5. Enterprise AI’s bubble will break at the renewal date
Riparbelli agreed with executives still waiting for measurable AI ROI. Many buyers have been instructed to produce an AI strategy, making them eager to hold conversations and fund pilots from innovation budgets, yet they lack enough technical understanding to identify what their businesses actually need.
That creates a deceptive signal for vendors. A buyer who has just spent $100,000 may insist the project works, while a startup interprets payment and praise as delivered value; the mismatch surfaces only when 12-month agreements expire and “that wall of churn” arrives.
Consumers will pay $30 monthly to try something novel, and enterprises will sign $50,000 pilots to demonstrate activity to management. Riparbelli’s dividing line is uncompromising: “The real signal is not that you sign a contract. The real signal is renewal.”
He nevertheless called the bubble a productive feature of capitalism: money funds millions of experiments, most fail, and useful products emerge. He expects substantial money to go up in flames on products that are not valuable or eventually become features of major cloud providers. His yellow flags are technology-first buzzwords such as “AI agents” and especially “AI employees”; these are algorithms and software, not colleagues, and anthropomorphism distracts from the customer’s problem.
6. Distribution matters more as foundation models converge
Riparbelli expects strong engineers to become substantially more productive over the next six, 12 or 18 months, but rejected claims that software businesses reduce to generated lines of code. Prompting a browser Tetris demo is far removed from operating complex products shaped by customers, maintenance, feedback loops and human organizations.
Stebbings sharpened the objection: a company maintaining 150 internal tools would also inherit responsibility for continuously upgrading 150 tools. Riparbelli called the idea that anyone can request a Monday.com clone and destroy the incumbent at one-tenth the price “a really naive view” of building a business.
In text generation, current models are already capable enough for many use cases where LLMs will transform the world, so “distribution is king and great products are king.” OpenAI has captured the consumer destination, but Riparbelli sees model capability increasingly embedded across many applications rather than monopolized by one interface.
Given $10 million to invest at OpenAI’s stated $160 billion, Anthropic’s $60 billion or xAI’s $50 billion valuation, Riparbelli chose xAI. The reasoning was asymmetric upside: Elon Musk’s execution, X’s hundreds of millions of users and the platform’s real-time information could combine distribution, data and model access.
7. Scaling continues, but linear compute forecasts will fail
Riparbelli does expect further gains from scaling, especially beyond text in video, audio and 3D. He does not expect a simple law where whoever owns the most compute wins: “The world rarely works that way.”
His three inputs are compute, algorithms and data. An algorithm that becomes 10 or 100 times more efficient could reset the competitive order, so extrapolating today’s training budgets misses the possibility of discontinuous technical improvements.
GPT-5 will probably be better and excite specialists tracking benchmark gains, but Riparbelli doubts it can recreate ChatGPT’s original cultural moment. Humans “saturate incredibly easily,” and the capability required to astonish users rises after every release.
His time-travel test captures that shifting baseline: ChatGPT shown 200 years ago might get its operator burned at the stake; 50 years ago it would unquestionably be called AI; today it is dismissed as “just a stochastic parrot.” For most businesses, he argued, existing models already solve many practical problems.
8. AI’s difficult frontier is control over outputs and use
Video generators such as Sora and Runway can produce striking material from almost any prompt, but Riparbelli compared the experience to a slot machine: enter a request, receive an output, adjust the prompt and “pull the slot machine again.” That is impressive until the user needs something exact.
Commercial control means preserving the same character across scenes, delivering specified dialogue and changing individual elements predictably. Each constraint can reduce overall fidelity because the model must follow instructions instead of merely reproducing something plausibly real; better control will require algorithmic advances.
Riparbelli described roughly 99.9% of content as obviously acceptable “green” material and recognized a clear red zone around hate and violence; the hard territory is gray, such as distinguishing an enthusiastic cryptocurrency founder from an outright fraud promising a 10x return.
Riparbelli supports community-driven systems resembling Wikipedia and Community Notes over centralized human judgment, while conceding they remain unsolved. Synthesia is stricter: enterprise customers do not want branded avatars beside conspiracies, and collecting $30 monthly from questionable creators is poor economics, so “today we are” arbiters of acceptable use.
9. Workflow, not avatar generation, protects the application layer
Riparbelli called predictions that OpenAI would simply move into Synthesia’s category “classical VC brain”—an excessive focus on the underlying technology. Customers do not principally buy an avatar model; they buy a workflow for conveying a message efficiently and engagingly.
That workflow spans drafting, camera-free presentation, voiceover, a PowerPoint- or Canva-like editor, multilingual playback, hosting, publishing and distribution. Million-dollar-plus contracts reflect the complete value chain, while avatars remain the conspicuous headline feature rather than the whole purchase.
Synthesia builds proprietary models only where it believes it can be the world’s best, the domain is narrow and the capability compounds inside the broader workflow. Its chosen niche is dialogue-driven video: voiceovers and humans presenting to camera.
Elsewhere, the company mostly uses OpenAI, while shifting some workloads to Anthropic and also using Gemini. Price matters for large workloads such as moderation, alongside task-specific quality and scaffolding; many future “specialized models” may simply be tuned open-source components operating invisibly inside useful products.
10. Production costs can vanish without eliminating creative skill
The internet democratized distribution, and smartphones partially democratized capture; generative systems now move creation from recording the physical world with sensors to producing media digitally. Riparbelli compared the transition with software instruments, which let a laptop recreate music without physical instruments.
From a purely technical perspective, he expects the cost of producing a Hollywood film eventually to approach zero. The result is “an absolute deluge of content” and a flattened production-value gap: YouTubers and independent creators can compete with studios on what they bring to life.
Stebbings’ pushback was grounded in current production: automated clipping tools are “incomparable” with his team, and the best extract from an episode depends on the intended outcome—six different sections could each legitimately become the highlight.
Riparbelli conceded the distinction between removing technical barriers and removing skill. Cameras, distribution relationships and eight years of visual-effects training may matter less, but storytelling, cultural judgment, relatability and asking the right questions matter more. He also expects AI to make a first pass that is better than random but worse than expert work; an expert plus AI could be extremely effective. Human-made work may command more value, much as audiences still value watching a real pianist.
11. Discovery and provenance become scarcer than content
Asked whether five years brings more human- or AI-made material, Riparbelli answered “definitely AI-made content.” The decisive question is what surfaces: roughly 95% of YouTube videos, he estimated, have fewer than 10 views, so cheap supply does not itself create demand.
TikTok’s advantage is an interest graph rather than a social graph. Riparbelli described his feed as roughly 90% educational material about technology, politics and music; the platform tests a new post on a few hundred viewers, discards weak responses and expands distribution when engagement appears.
Both speakers resisted treating social media as inherently poisonous. Stebbings compared feeds with a food diet, while Riparbelli wanted more explicit algorithmic controls. His provocative possibility: a two-minute TikTok might sometimes explain a subject better than a 20-minute video or 300-page book, rather than merely serving “dopamine junkies.”
For trust, Riparbelli proposed universal identity and content provenance: a Shazam-like fingerprint could identify who created something, when, with which tools and whether the viewer sees an edit. He described a centralized database—not ideally a decentralized registry—while adding that it could actually be a blockchain. Such a system might expose a five-year-old war image falsely presented as yesterday’s event.
12. People and policy—not models—set the adoption curve
Creative roles will change unevenly: some camera operation may fade, while visual-effects artists using AI could become “100 times as efficient.” AI may handle more of the technical first pass, but skilled creators still determine what matters and how it serves the audience.
Riparbelli expects far more people to become creators, just as computers turned text production from a specialist activity into something everyone does. The principal rate limiter is already human: people must trust AI, purchase it, integrate it and alter established working habits.
His unusual training thesis is that complex games teach decision-making and potentially entrepreneurship. Warcraft 3, World of Warcraft, Red Alert and RollerCoaster Tycoon create rapid decision-feedback loops; even a simulation that captures only 20% of running a real theme park provides iterations unavailable before computers. “Life is a game. Careers are a game. Building companies is a game.”
Riparbelli also reframed his lack of specialization as an advantage. Curiosity across programming, music, elevators and Norwegian black metal trained a broad “neural network” for analogies and pattern matching; he now sees being decent at many things as useful for a CEO whose work is largely decisions and ideas.
13. London offers loyal talent, but policy can squander it
Riparbelli’s honest reason for founding in London was that he could not obtain a US visa or sponsorship. He still believes the probability of startup success is higher on the West Coast, but says London has improved markedly over six or seven years and can draw Americans attracted to its cultural breadth and European lifestyle.
Europe’s underappreciated advantage is loyalty. Silicon Valley employees often rationally build portfolios of startup options and leave after bad quarters; Europeans still treat options more like lottery tickets because few know someone who made $2 million at a startup, producing a less transactional relationship with employers.
The UK now has companies capable of global rather than merely regional leadership, including ElevenLabs in voice, Synthesia in enterprise AI video and Wayve among the top three self-driving companies globally, in Riparbelli’s view. He urged government to protect or improve entrepreneurial relief, avoid excessive taxes and regulation, and consider subsidizing GPU costs through UK data centers.
His warning was blunt: “Europe has all the AI regulation and none of the AI companies.” He called the London Stock Exchange’s lack of liquidity “a disaster,” though not a top-three policy priority. He argued that one or two billion-dollar-plus companies headquartered in the UK could help the ecosystem grow; Stebbings suggested a couple of $10 billion companies, and Riparbelli agreed.