ElevenLabs CEO: Why Voice is the Next AI Interface
Summary
- ElevenLabs keeps research ambitious without letting it become a product bottleneck. After resisting a simple speed control for nine months because the company hoped research would make voices infer pacing automatically, it adopted a three-month rule: beyond that horizon, product teams may bridge the gap however they choose. “We don’t want to become the same as the previous generation of the editing suite” was the rationale for resisting the slider, not an absolute constraint.
- Shipping velocity comes from roughly 20 autonomous product teams of five to 10 people, backed by a research foundation. That structure tolerates duplicated work and uneven speeds in exchange for unusually high ownership and the ability to pursue creative tools and conversational agents simultaneously. Each new team gets six months to prove itself; the infrastructure team grew from three people at the partnership to 11.
- Its distributed model is a global-talent thesis, not merely a remote-work policy. ElevenLabs hired globally and unconventionally—including an open-source text-to-speech developer who was working in a call center—then added hubs once headcount exceeded 30. Titles were removed, tenure does not determine hierarchy, and information access is deliberately limited when transparency becomes distraction.
- The voice marketplace converts model breadth into an ecosystem with measurable creator economics. Nearly 10,000 voices are available and $10 million has been paid back to contributors; one deep Spanish voice found little demand in Spain but became a top-three voice after being offered in other languages and taking off in an English-speaking market because of its deepness. Staniszewski’s preferred posture is to help industry participants “disrupt together rather than just disrupt.”
- Enterprise expansion required abandoning the conceit that engineers could simply perform sales. The customer-facing mix is now roughly 80% sales and 20% engineering, while the product has expanded from text-to-speech into speech-to-text and orchestration, and enterprise deployments require telephony, evaluation, monitoring, security, and compliance. The production ambition is eventually “four nines or five nines,” though Staniszewski concedes that reliability is difficult in AI.
- At 350 employees, incentives have become part of product and competitive strategy. Staniszewski calls quotas and commissions “a lagging indicator of strategy”: the company may still grant commission while killing a strategically wrong deal, including a recent request from a foundation-model competitor to license ElevenLabs models for demos. The resulting policy explicitly prohibits sales to foundation-model companies.
Deep dive
1. Research and product run on deliberately different clocks
Staniszewski credits ElevenLabs’ foundation to research that made text-to-speech understand context, translate it into “emotion and intonation,” and preserve voice characteristics such as style, age, gender, and dialect. That base subsequently expanded into speech-to-text, music, and other audio work. He also contrasted a three-person infrastructure team at the time of the partnership with 11 people now.
Roughly 20 product teams, each five to 10 people, can ship independently across two broad domains: creative production—narration, voiceovers, and dubbing—and conversational agents spanning customer experience to immersive media. The accepted costs are duplicated work and teams moving at different speeds; the payoff is ownership.
The instructive failure was a requested voice-speed slider. ElevenLabs resisted it for nine months, hoping research would make each voice choose the right pace automatically, before conceding that the simple product fix served users. The resulting rule: if research will take more than three months, product teams can add other models or extensions to close the gap.
2. Global talent and flat teams are operating choices
The company’s European origin shaped both product and hiring. In Poland, foreign films were often narrated by one emotionless voice, regardless of character; addressing that research problem required researchers across Europe and Asia rather than restricting recruitment to San Francisco.
Unconventional sourcing produced an open-source text-to-speech developer who was simultaneously earning money by taking call-center calls. After headcount passed 30, ElevenLabs added hubs in London, Warsaw, and San Francisco: early-career hires are generally placed in hubs for immersion, while experienced remote workers retain flexibility.
Titles were removed a year earlier, and new teams receive six months to prove themselves. Tenure does not fix anyone’s place in the hierarchy, but flatness still needs leads who can carry context across teams. One counterintuitive lesson: putting everyone in every Slack channel created distraction, so access is sometimes cut “to force the attention.”
3. Creator alignment turns voice supply into a moat
Staniszewski argues that adoption starts with learning which production steps creatives want AI to touch. ElevenLabs’ marketplace then lets people share voices and earn money in return when a voice is shared: it now contains almost 10,000 voices and has returned $10 million to the community.
His best example is a deep Spanish voice that initially failed to gain traction in Spain. Because the same voice could speak 30 languages then—and 70 now—it took off in an English-speaking market because of its deepness and became a top-three voice across use cases.
Music licensing applied the same “disrupt together” philosophy. Reaching an agreement with Merlin, Kobalt, and four major labels took 18 months, with forcing timelines moved a few times but still useful for creating urgency around whether to proceed together or separately. The licensed model can grant commercial rights to generated music.
Industry-native advisers helped bridge unfamiliar negotiations, but Staniszewski found that risk judgment mattered as much as background. A lawyer who had worked at Fortune 500 companies framed nearly every proposal as a list of risks; the replacement counsel explains the risk boundary, comparable-company behavior, and a workable course of action—a “true thought partner.”
4. Enterprise demand pulled ElevenLabs from models into infrastructure
ElevenLabs initially wanted an engineering company without salespeople and tested one conventional seller against an engineer told to “do sales.” It failed. The eventual formula—about 80% sales and 20% engineering—gave research and product teams a clearer view of customer requirements.
Hippocratic AI supplied the pivotal healthcare example in 2023: voice agents could take inbound hospital calls, schedule appointments, remind patients about medicine, and conduct outbound follow-ups. Delivering that required speech-to-text, an LLM, text-to-speech, orchestration, integrations, and deployment—not merely one foundational model.
The larger enterprise gap is between a demo and production: testing, version control, evaluation, monitoring, and fine-tuning over time. Closing it requires orchestration and surrounding deployment capabilities, including knowledge bases, telephony providers such as Twilio and SIP trunking, security, and compliance; dependable “four nines or five nines” remains an aspirational endpoint.
Li’s pushback is that enterprise demands often slow launches. Staniszewski’s answer is explicit labeling: customers choose whether to access alpha products that “might not be as stable.” Deutsche Telekom, for example, is creating new podcast experiences after testing early models for NotebookLM-style podcasts in selectable German and English voices.
5. Scaling makes portfolio rules and incentives unavoidable
Once ElevenLabs exceeded 100 people, it separated pre-product-market-fit work from mature products. Mature teams test for the long term and deploy only when ready; experimental teams have six months to find product-market fit through rapid shipping. If they cannot prove a substantial user base, “we kill the product.”
At 350 employees, Staniszewski’s sharpest change of mind concerns incentives: passion no longer coordinates every decision, and commissions can produce behavior strategy never intended. The company may grant commission while leadership rejects a strategically harmful deal; after declining a foundation-model competitor’s licensing request, ElevenLabs formalized a ban on selling to that category.