Building One of AI’s Fastest-Growing Companies | Mati Staniszewski, ElevenLabs
Building One of AI’s Fastest-Growing Companies | Mati Staniszewski, ElevenLabs
Summary
- Senra likened ElevenLabs to Honda’s engine lab, and Staniszewski agreed: it is a focused audio research lab with a product platform layered on top. His filter for every new product is: “do we believe we have a unique advantage by applying our audio models to that product experience?” That is why they paused avatars two years ago (video was the bottleneck, not audio) and explicitly will not touch “intelligence, knowledge work, or programming,” calling that “a very, very fierce battle.”
- The company started in 2022—before ChatGPT, when “the topics of the day were cryptocurrencies and the metaverse”—to fix Polish single-voice movie dubbing, then detoured when creators said dubbing was not their most urgent problem. All three dubbing steps (transcription, translation, regeneration) were “quite robotic” in 2022, and creators asked instead for post-production fixes and AI narration. A model able to carry content from one language to another launched “just a week ago.” The guiding star was Hitchhiker’s Babel fish—not building the fish, but “allow[ing] everyone to have the Babel fish in their devices.”
- Growth is agents-led, and the expansion playbook is Palantir-style forward-deployed engineers who sit inside the product team, not marketing. Deutsche Telekom went from a marketing podcast two years ago to call-center voice agents to an in-network agent that joins T-Mobile calls to translate in real time. Fintech is the fastest-adopting vertical (Revolut, Klarna, Nubank), followed by healthcare and telecom, with retail “only just got started this year”; the largest revenue lines are agents and creators.
- Staniszewski concedes the pure-research edge over big labs “might not be as significant” long-term—the hedge is ecosystem: a 20,000-voice authenticated marketplace where voice owners earn passively. The “George” voice used in David Senra’s ElevenReader research reports pays a real voice actor on every listen. His stated aspiration over the next three years is a world with “3 platforms where you centralize all interactions”; ElevenLabs wants to be the leading platform for that communication.
- ElevenLabs has turned down three or four concrete acquisition offers, the last around June of last year, with none active. “AI is changing the world. We can build the frontier of that change. We’re going all in.” Senra pressed hard: at 31, ElevenLabs “is probably the best idea you’ll ever have”—why spend four decades on your fifth-best? Staniszewski: “you’re trying to convince me of something I’m already convinced of.”
- The 12-month call: AI conversation that connects IQ with emotional intelligence—understands how you feel, pauses, re-enters the conversation—flipping decades of humans learning “the language of technology” to “bring the technology to our terms.” Senra’s analogy: telegraph versus telephone—“you just pick it up and do exactly what you already do.” Counterintuitively, making agents imperfect (“uh,” “ah,” pauses) made performance “skyrocket.”
- Org design as thesis: sub-10-person teams, a hard cap of five management levels (fewer over time “if AI helps us run the organization”), engineers embedded in legal, talent, and ops, and broad document transparency. The AI SDR on the website—an alternative to the dropdown form—gets prospects to “leave much more information than they would ever leave in a drop-down form.”
Deep dive
1. Founded before ChatGPT, out of a Polish dubbing grievance
- Staniszewski and co-founder Piotr—“my best friend of 15 years… the smartest person I know”—started ElevenLabs in 2022, pre-ChatGPT, when “the topics of the day were cryptocurrencies and the metaverse. So it was a perfect time because we could focus and build a lot in the field of AI.”
- The trigger: in Poland, every film is narrated by one voice for all characters—“All emotion and intonation disappears”—and in 2021 it was exactly as it had been in his childhood. The founding vision was original voice, original emotion, in your language; a model able to carry content from one language to another “extremely well” launched just a week before this recording.
- The résumés fit the research-plus-deployment thesis: Mati built optimization models at Palantir (NHS COVID vaccine distribution across the UK, oil-and-gas energy work) after risk models at BlackRock, with a math degree; Piotr was in charge of many text models for Google’s Knowledge Graph.
- The Hitchhiker’s Guide Babel fish was literally “one of the slides”—but the goal was never to build the fish itself: “we would allow everyone to have the Babel fish in their devices, their presence, and their current work.”
2. Dubbing was the idea; the market forced a detour through voice generation
- Dubbing decomposes into transcription, translation, and regeneration—and in 2022 the research was not there: “everything was quite robotic,” and nothing had crossed the uncanny valley. NVIDIA had good models and open-source components, DeepL handled translation “incredibly well,” and the best open-source repository was “quite good but very unstable.”
- Early creators delivered the pivot: “Great, I would love to have dubbing someday, but today I have other problems”—fixing a misrecorded line in post, hearing a script aloud before filming, or having AI narrate a video. So the plan became: solve voice-generation research first, ship narration, and “ignore the language change for a second.”
- Senra’s corroborating anecdote: MrBeast told him years ago he did not just dub his Japanese videos—he hired voice actors already famous in that country.
3. Run it like Honda’s engine lab: research is exclusively audio
- Senra’s framing from his Honda episode: Soichiro Honda called his company “just an engine research lab,” spun R&D into an independent company financed by a percentage of sales, and rejected any product without a motor. Staniszewski’s response: “this isn’t too different from how I think about how we run ElevenLabs today”—a focused research lab, research engineering, and “many small teams, usually of fewer than 10 people,” with autonomy to apply their best judgment.
- Why audio-only: of data, compute, and architecture, “many of the problems that still exist are on the architectural side” of audio—and audio is “a good combination not only of science, but also of the art,” because voices are subjective. Asked whether they have technology no one else in the world has: “We believe so.”
- The Honda-style filter is asked internally on every new product: “do we believe we have a unique advantage by applying our audio models to that product experience? If the product experience doesn’t have a major bottleneck on the audio and voice-communication side, then it’s not our strength.”
4. What they said no to—and the ecosystem hedge against the big labs
- The best specimen of discipline: two years ago they paused avatars and lip-sync because “the problem to solve there is the video, not the audio”—the best audio could not rescue bad video. Today, open-source video models are good enough that they are picking it back up, but “the conscious decision we’re making is still not to create the model from scratch.”
- Equally explicit: “we explicitly don’t plan to touch anything related to intelligence, knowledge work, or programming. It is not our strength… A very, very fierce battle.”
- Is the audio focus a defense against larger labs? “Definitely”—but he concedes the pure-research advantage “might not be as significant” in the very long term, which is why product and ecosystem matter: a marketplace of 20,000 authenticated voices where creators earn passively as their voice is used. The “George” voice used in Senra’s ElevenReader research reports pays a real voice actor on every listen.
- The stated aspiration for the next three years is “3 platforms where you centralize all interactions”; ElevenLabs wants to be the leading platform for those interactions and that communication.
5. Deutsche Telekom is the land-and-expand template
- Entry was marketing two years ago—agents then “weren’t very reliable, they weren’t very fast”—with a daily AI-voiced podcast in the Magenta app plus ads. Then came call-center voice agents for support and billing; most recently, an in-network agent a T-Mobile subscriber can ask to join a call to schedule a reservation or translate the conversation in real time.
- The expansion mechanics: “we’re not just testing concepts; we test the impact, we test the value” before scaling—then the unglamorous work of CRM integrations, SIP-trunk or Twilio phone connections, and the hardest part, verifying agent behavior through simulations and call testing. Forty-eight full-time engineers partnered side by side with the team in Germany, and “the work doesn’t end—you still want to evaluate, monitor and refine over time.”
- Vertical adoption ranking, as stated: fintech fastest (“Revolut, Klarna, Nubank… moving at another speed”), healthcare and telecom next, and retail “only just got started this year.” Revenue-wise, “the largest lines are on the agent’s side and the creator’s side.”
6. Palantir’s fingerprints: FDEs live in the product team
- The formative contrast: at BlackRock, his first month or two meant outbound emails were verified by his team; at Palantir, weeks after onboarding, “you’re going to fly to Aberdeen, on the North Sea, work side by side with them”—“a total shock to me… I thought it was great.”
- He also valued a pattern he saw at other companies: small frontline deployment teams, around five people at most, where “the best-idea-wins” approach gave the team power to act.
- His sharpest structural point: FDEs at ElevenLabs sit in the product team, not marketing, because the second, “frequently overlooked” job is bringing knowledge back into the product—otherwise “it’s basically just services.” Senra’s summary, confirmed: the customer base becomes another form of R&D, since “rarely are a company’s problems unique to that company.”
- The stranger Palantir inheritance: the internal “artists’ colony” concept—“never said publicly”—which he now values more as AI meets creative work. On taste-as-buzzword: “beyond the buzzword, I think there’s a lot of truth in that… more and more everyone will be able to create everything,” so design language and refinement decide who wins. Senra’s addition, via Edwin Land: “good taste is as rare as a unicorn.”
7. Org design for the AI era: flat, small, engineer-saturated
- Palantir had no job titles and a relatively flat organization. At ElevenLabs, most people have close to 10 direct reports, and there is a hard cap of five management levels—“hopefully, over time, there will be fewer, not more, levels if AI helps us run the organization the way we’d like.” Small independent teams also mean AI adoption happens bottom-up, not by mandate.
- What AI changes in management: leaders get signal “from the person who works closest to the problem” instead of manager summaries, and leadership shifts “from reactive to proactive.” The prerequisite is a deliberate risk: “almost everyone has access to all the company’s documents.”
- Every team—talent, operations, and legal—has engineering talent to automate work and help the team use AI. Senra’s comparison: Bending Spoons’ Luca Ferrari has an HR department of 50 people “and they’re all engineers.” Staniszewski expects this pattern to generalize across companies.
- The AI SDR is the concrete payoff: visitors can talk to a voice agent instead of filling a dropdown form—they enjoy it more and “leave much more information than they would ever leave in a drop-down form.” It is developed centrally and refined by each local team.
8. Voice’s moat is emotional—and imperfection is the feature
- The accessibility record: over 10,000 people have had voices restored (including people with ALS or throat cancer), plus partnerships with 800 organizations—a woman who lost her voice before her wedding and could say her vows in her recreated voice, a musician now touring UK venues with his AI voice, and a congresswoman appearing in Congress with hers. The showcase: Tim Green, an NFL player and former best-selling author with ALS, launched a podcast with his AI voice and won an Emmy this year—“hundreds or thousands of people contacted us because of him.”
- The counterintuitive engineering lesson: they first tried to build a flawless agent—“It didn’t sound human.” Adding “uh,” “ah,” and pauses made performance “skyrocket”: “making it imperfect is almost the key element of it now.”
- Why audio research stays hard: every voice is different and subjective—even benchmarking text-to-speech is “extremely difficult, because different models typically have different voices. That already makes them incomparable.”
- On Senra’s nationalization question, an honest non-answer worth keeping: “I don’t have a direct answer”—but they spend real time with European teams and governments on sovereignty, and his conviction is that AI “needs to amplify human potential rather than replace it.”
9. Three or four acquisition offers refused—“we’re going all in”
- Concrete offers: “three or four,” the last around June of last year, none active. His full answer: “AI is changing the world. We can build the frontier of that change. We’re going all in.” Any of the acquisitions “would have made life very easy, of course. But I don’t think money alone should be an interesting proposition.”
- Senra’s extended pushback is the episode’s emotional core: founders who sold tell him “I’d pay back the billions if I could get my company back”; at 31, “chances are ElevenLabs is probably the best idea you’ll ever have… you’re going to work on your second, third, fourth, fifth best idea?” His Steve Jobs test: no sum could have kept Jobs from Apple—“if you love what you do, they couldn’t pay you to do it.”
- Staniszewski, gently deflecting the sermon: “you’re trying to convince me of something I’m already convinced of.” Senra separately argues that Scott Wu’s Cognition should also build an independent company. Staniszewski says co-founder Piotr is “totally committed” as well; “we all realize how unique this moment is.”
10. The next interface is voice—the 12-month prediction
- The bottleneck thesis: “Intelligence, in general, is still evolving. The next bottleneck in accessing that intelligence will be how you communicate and collaborate with it.” In 12 months: AI that connects “the IQ aspect with emotional intelligence”—knows how you feel, pauses, thinks, and re-enters the conversation. “For decades we’ve learned the language of technology, the keyboard, the screen… now we can bring the technology to our terms”—ideally the screen and phone “can stay in your back pocket.”
- Senra’s historical rhyme: Alexander Graham Bell on telegraph versus telephone—the telegraph required learning a language; the telephone, “you just pick it up and do exactly what you already do.” If voice enables that transition, “the number of people using AI will explode dramatically.”
- A closing operator’s note on conferences: Staniszewski does two to three per quarter, but only after learning the hard way—his first, unprepared, “was terrible.” Now it is pre-scheduled one-on-ones with partners and customers, never just the sessions. Senra’s counter-motto, from a friend: “Stay away from the circus.”