a16z on AI Voices: Call Centers, Coaches, and Companions with Olivia Moore & Anish Acharya
Summary
Voice AI’s first durable commercial wedge is vertical B2B, not a standalone consumer app. Businesses already pay people to answer phones, making after-hours coverage and otherwise unanswered calls easy entry points; the strongest startups then expand into workflow ownership. Labenz says some are among “the fastest-growing B2B startups we’ve seen in 10 years.”
Conversational viability is largely here, but human likeness still depends on more than transcription accuracy. Latency is now typically below half a second, while Sesame showed how pauses, “ums,” and vocal inflection can turn a polished synthetic voice into one “that could be mistaken for a human.” Emotional adaptation, multi-party turn-taking, and interruptibility remain material gaps.
Happy Robot demonstrates why specialized conversation quality can unlock higher-value work. Its agents disclose that they are AI and befriend, disagree with, and negotiate with truckers for freight brokers; one tactic inserts a five-second “let me talk to my supervisor” delay before returning with a concession, which Moore says, with some uncertainty, produces a much higher acceptance rate. The investable insight is that better delivery earns permission to handle persuasion and pricing, where commodity voice does not.
The application moat is the vertical system around the voice, not voice generation alone. Enterprises need integrations, customer-specific context, evaluation, guardrails, and recovery when systems of record fail—“the capability gets you in the conversation but isn’t sufficient to get you to the other side.” That favors vertical platforms over horizontal agents, even as base models improve.
Apple’s postponed Siri overhaul illustrates an incumbent disadvantage, not a lack of technical progress. Acharya calls Siri “a stick in the eye five times a day” and argues that large companies struggle to embrace AI’s messy humanity; Moore adds that Apple must ship safely to hundreds of millions of users, unlike a startup serving self-selected beta testers. Labenz presents Google’s failure to commercialize Deep Research before ChatGPT became associated with it as a parallel missed opportunity.
Voice automation is producing coaching and task substitution, but not yet the 90% call-center headcount collapse Labenz tested as a scenario. Real-time coaching can justify hundreds of dollars per month when it influences a $10,000 HVAC upsell, while recruiting agents can return roughly 20 hours a week for work with five priority candidates. Despite call-center turnover reaching 300% annually, the guests had not observed order-of-magnitude job losses, and Acharya would not confidently translate an 18-month technology horizon into labor-market timing.
The consumer frontier extends from seniors and children to increasingly personalized companions. Multimodal voice could give seniors patient technical help, provide children with tutors or socially positive Minecraft partners, and produce companions ranging from sympathetic listeners to challenging “East Coast mode” personalities. Acharya’s longer-term framing is an “emotional bicycle” that extends people emotionally as computers extended them intellectually.
Safety policy must reconcile demonstrated impersonation risk with a market for licensed identity. Labenz reported that two calling platforms still let him clone Donald Trump and make scalable calls a year after he disclosed the problem, while Moore said people were currently more frustrated by model restrictions than by being cloned. Her twist on a do-not-clone registry is economically constructive: let people prohibit impersonation while explicitly licensing their voices or avatars for approved uses.
Deep dive
1. Demand-side behavior is often the clearest product signal
Moore’s scouting funnel spans founder-heavy Twitter, newsletters, and meetups, but also Instagram, TikTok, and especially YouTube—the “number one mobile app and the number two website in the world.” For many consumer and prosumer AI products, YouTube tutorials are the largest source of social referrals.
The team watches ordinary users—“usually teenage girls, to be frank”—force ChatGPT into roles such as therapist, friend, or coach. Moore’s framing: “Consumer is so random and magical that we try to let the data tell us,” because pedigree cannot rescue a consumer product that misses timing, onboarding, or one decisive feature.
Moore’s less glamorous edge is simply using the products: Operator, Deep Research, DeepSeek, o1 Pro, and Krea were examples of tools many supposed insiders had not tried. Her memorable demand test updates an old social-app joke: “Every large language model is being tortured into being a therapist.”
2. Voice opens a technology surface that screens never addressed
Acharya starts from first principles: voice intermediates most human relationships, yet technology historically lacked the infrastructure to address it. Unlike computing surfaces with decades of product experimentation, voice remains “a complete blank piece of paper,” creating both product and distribution opportunities.
Moore sees more traction in B2B agents because even small businesses pay one to three people to answer phones. Once models approach human performance, using an agent for after-hours calls or calls otherwise sent to voicemail becomes an obvious substitution; many consumers may already have encountered one unknowingly.
Consumer exposure has instead flowed through ChatGPT, Grok, and the viral Sesame demo. Moore reads 1-800-ChatGPT as a clue that many people’s first meaningful AI experience—whether or not that product itself succeeded—may arrive by telephone rather than through a new app.
3. Seniors show why multimodality matters as much as conversation
Moore’s early-90s mother already asks Alexa to play music, making Alexa+ a plausible bridge into conversational computing. Voice could also unlock older technologies she never learned to operate, from email interfaces to the television remote.
Labenz’s reliable support instruction—“read everything on the screen from the top to the bottom”—already worked as a text-only GPT-4 prompt. A patient model could reproduce that help for people who lack an “infinitely patient” relative.
Acharya points to Google’s release—“I think it was late last year, in December”—of Gemini models that could see what was on a screen and interact in real time; OpenAI has a similar capability. His sharper point is that the relevant context is often physical: pointing a phone at the remote is more natural than verbally translating every button to Nathan or Google.
4. Sub-half-second latency solved entry-level voice, not conversation
Acharya says “solved” is too strong, but basic latency and understandability moved close to solved within the past year. Most models now respond in less than half a second—the difference between exchanging audio and being able to sustain a conversation at all.
Sesame’s advance was to preserve apparent imperfections: extra pauses, fillers such as “um,” and expressive inflections that a conventional system might treat as errors. Those details move output beyond a better Alexa or Siri toward a voice “that could be mistaken for a human.”
Emotionality remains unfinished. Founders want agents to recognize content and adapt tone—brighter for exciting news, lower and slower for sadness—while interruptibility remains awkward, especially in groups. Moore notes that humans themselves have not fully solved the moment when two people begin speaking together.
Labenz distinguishes voice models from conversation models: people coordinate turns through subtle audio and visual signals and often begin speaking before knowing the exact sentence. An AI that already knows its entire response can feel uncanny, suggesting that native conversation behavior needs more than a speech-to-text-to-LLM-to-speech pipeline.
5. Specialized delivery earns the right to negotiate
A16z portfolio company Happy Robot supplies voice AI to freight brokers. Acharya argues that its deeper technical work makes conversations feel materially better than commodity alternatives, which gives customers confidence to delegate “persuasion, negotiation, disagreement”—not merely information retrieval.
Moore’s standout example is deliberately added latency. The agent says, “Hold on, let me go talk to my supervisor,” waits about five seconds, then returns with a slightly better price, despite already knowing its permitted range; the human feels they earned a concession, and Moore thinks the final-offer acceptance rate is much higher.
Labenz’s reaction—“I don’t know how to feel about that”—preserves the tension between effective design and manipulation. The agent discloses that it is AI, yet truckers fall into familiar conversational rhythms anyway; intellectual awareness does not override the “reptilian brain” trained by negotiation rituals.
The broader claim is that AI can be “more human than the humans”: the same friendly agent answers each time, listens patiently, and can spend all the time in the world with the caller. Superhuman patience and low-to-no wait times are operationally valuable even before the model becomes superhuman at reasoning.
6. Incumbents face a cultural and distribution handicap
Apple’s announcement that a major Siri update would wait until 2027 feels absurd against rapidly improving startups. Acharya says the mismatch between poor everyday Siri performance and Apple Intelligence advertising degrades trust: “It’s like a stick in the eye five times a day.”
His diagnosis is institutional: AI works best across a messy surface of human interaction, while incumbents are organized to remove humanity and unpredictability from technology. Committees, lawyers, and efforts to “neuter the AI” create what he calls a “very difficult spiritual problem.”
Moore offers the counterweight: Apple must deliver something natural and correct to hundreds of millions of users across ages and use cases. Startups can ship imperfect beta products to voluntary early adopters; she guesses that public reaction to generated notification summaries spooked Apple.
Google Labs offers a partial model through gated experiments such as NotebookLM, but commercialization remains slow. Labenz’s sharpest example is Deep Research: it originated as a Gemini product that Google should have dominated, yet ChatGPT became known for the capability—“one missed opportunity after another with incumbents.”
7. Vertical context and integrations form the real enterprise moat
Native voice-to-voice models are earlier and somewhat more expensive; Acharya calls Gemini Flash perhaps the best current option, while noting that interruptibility still lags. His standing hedge is that “the models are the worst that they’re ever going to be right now,” so today’s multi-stage stack may look antiquated within a year.
Acharya sees reasoning models as a separate primitive despite their familiar interface. Probabilistic language behavior helps with friendship and rapport, while pricing and other factual decisions demand accuracy; orchestrating reasoning and language models can allocate each part of the conversation to the appropriate system.
Tool use makes voice generation only the opening layer. As Acharya puts it, “The capability gets you in the conversation but isn’t sufficient to get you to the other side”; durable products require workflows, integrations, outcome measurement, guardrails, and long-tail handling of backend failures.
Moore says this burden is especially severe for traditional enterprises that may struggle to build once, let alone keep models and systems of record synchronized. She emphasizes that vertically focused platforms can handle the long tail of industry conversation types and integrations; Labenz frames customer-specific context as the scarce “last mile” that a base model cannot infer from generic knowledge.
8. Coaching works now, while mass displacement remains unproven
Acharya says real-time coaching is already working for call-center staff, salespeople, and HVAC technicians. If one nuanced question determines a $10,000 upsell, even hundreds of dollars per month for an individual coach can deliver obvious returns; jobs with physical or deeply personal components remain particularly suited to augmentation.
Automation is also removing undesirable tasks. Call centers can experience roughly 300% annual turnover, while recruiting agents can conduct initial screens and return about 20 hours weekly for a recruiter to spend with five preferred candidates, including persuasion and candidate care.
Labenz rejects an easy reskilling story: displaced call-center workers will not necessarily move into higher-value jobs at the same employer. He asks whether the technology could support a 90% headcount cut and suggests that such an effect might be a 2026-plus phenomenon, while acknowledging that the question is about order of magnitude rather than a literal zero-human call center.
Acharya pushes back empirically: they had not seen such reductions, because jobs bundle screening with interviews, salary negotiation, onboarding, and even taking an employee to a baseball game. “Even if the technology is 18 months away,” he says, its labor-market effect remains hard to predict; abundance may shift the question from jobs toward purpose.
9. Consumer voice expands from SMB receptionists to emotional infrastructure
For restaurants, spas, and home-services companies, Acharya recommends vertical rather than generic agents. Owners generally are not firing the core employee who answered phones; they redirect that person toward customer experience and growth while the agent handles missed and routine calls.
Creator tools already span ElevenLabs voice cloning and described-voice generation, likely Delphi-style interactive digital clones, and HeyGen avatars. Moore thinks fully synthetic podcast hosting may remain “a couple years away,” although she can imagine a world in which a scripted episode requires neither camera nor microphone.
Labenz cites Synthesis, Ello, and Super Teacher as examples of personalized tutoring, while Moore’s strongest example is a positive AI Minecraft companion replacing “toxic teenagers,” alongside a multimodal classroom observer giving parents and teachers feedback unavailable in under-resourced schools.
Companion demand has also defied assumptions: the guests see more AI-boyfriend and interactive-fiction behavior, often serving women, than a market dominated by “frisky young dudes.” Companions might supplement relationships, teach conversation or flirting, absorb emotional strain, or challenge rather than flatter users—the jokingly named “East Coast mode.”
10. Identity licensing could pair protection with economic opportunity
Labenz cites research from Stanford discussed in an earlier Replika episode: users reported reduced suicidal ideation and, more often than not, greater engagement with the outside world, though he questioned some of the data and noted that part of it was self-reported. His conclusion is conditional: beneficial companions exist, but predatory and addictive versions can also be built.
His red-team evidence makes impersonation concrete: two reasonably well-known calling platforms still allowed rapid Donald Trump voice cloning and scalable outbound calls a year after he reported the vulnerabilities. He therefore favors disclosure rules and a do-not-clone registry sooner rather than later.
Moore sees the registry’s positive-market counterpart: people could license approved uses of their identity. ElevenLabs’ voice collections already create opportunities for voice-over artists, while an influencer with 5,000 followers might monetize an extensible avatar even when a large brand would ignore the human creator.
Acharya warns against reflexive paternalism, arguing that consumers have learned not to trust everything in books, online, or on social media. Moore’s long-term vision is voice across AirPods, glasses, computers, and every product—sometimes two-way, sometimes transcription-only—while Acharya calls it an “emotional bicycle” that extends human emotional capacity.