Speechify's Weitzman on owning GPUs, ElevenLabs, and the AI talent war
Speechify's Weitzman on owning GPUs, ElevenLabs, and the AI talent war
Summary
- Cliff Weitzman’s core capex math: renting an H100 from a hyperscaler runs $35,000–50,000 a year versus about $30,000 to buy it outright — roughly 1.5× the purchase price. The hardware is warrantied for three years and, he imagines, may keep working for 10; owned, memory-co-located clusters also give Speechify cheaper training and inference. Speechify spends tens of millions on NVIDIA GPUs and pays six-figure premiums to skip delivery queues, potentially getting a year of Rubin access before hyperscaler customers.
- Weitzman says an NVIDIA deal with Blackstone, BlackRock, Apollo and Goldman Sachs — underwriting up to 25% of a GPU’s value as collateral — creates a liquid secondary market and a price floor, “exactly what Elon did in the beginning of SolarCity.” On circularity fears, he distinguishes the Oracle–OpenAI arrangements, which he calls ridiculous, from NVIDIA’s real, useful assets: “how many teraflops per second can this device do?” is effectively the unit of value.
- Weitzman’s confession is the episode’s centerpiece: not going B2B was “the biggest strategic mistake I made in the history of Speechify,” and ElevenLabs leapfrogging him was “100% on me.” He wrongly assumed a text-to-speech API would commoditize, missing that an AI lab’s first product is a wedge for everything after it; Speechify’s Simba 3.2 API now costs $10 per million characters versus ElevenLabs’ $100 and OpenAI’s benchmarked $196.
- Harry Stebbings argues going B2B now could be a fresh mistake — competing with an “unstoppable” ElevenLabs, which he says has Western-government buy-in, and Bret Taylor’s Sierra is “the Postmates effect.” Weitzman refuses to sit out: “The best way to lose is not to be in the race.” His precedents are Anthropic following OpenAI, Facebook following Friendster and MySpace, and OpenAI fumbling voice AI.
- On the AI talent war, Weitzman inverts Stebbings’ premise: hiring is brutal at growth stage, where packages can reach $15 million a year and Stebbings sees $50 million-plus, but “the easiest time ever” for true seed companies because raw aptitude can be taught rapidly. Speechify now hires math Olympiads, Kaggle winners and physicists who may never have coded: “hiring for slope more than intercept.”
- Speechify’s dev culture gives zero credit until code ships to production — “you make me a beautiful bottle of milk and you leave it down the road, the milk will spoil.” Claude Code is the top harness, engineers run 5–18 agents each and aim for “10 really good decisions per day.” There are no token leaderboards; wasteful token use can lead to people being let go, while a $12,500, two-week long-horizon run that produces a better model is money well spent.
- On public-market calls, Weitzman picks Meta over xAI, saying “Elon’s distracted.” He then compares Meta with Elon’s SpaceX and Tesla: Meta is around a 32 P/E while Tesla is valued in the multiple hundreds and SpaceX is “insane”; Meta has more data than anyone but is constrained by GDPR and other laws, and Zuck has roughly 20 extra years. Stebbings counters that removing Zuck could lift Meta’s stock by ending the “CapEx, CapEx, CapEx” focus, whereas removing Elon destroys much more value.
- His five-year contrarian call: human-computer interaction becomes primarily voice — and his deepest excitement is AI biology. For a family member with an orphan disease, he has analyzed 15 weeks of blood, genome, proteomic and RNA data against six years of self-reported data on a GPU cluster; he plans to sequence other patients to find a common thread. He says GPUs also helped identify his father’s prostate-cancer lesion, closing the loop on a founder story that began with dyslexia and Harry Potter audiobooks listened to 22 times.
Deep dive
1. The buy-don’t-rent GPU thesis: 1.5× per year to rent what you could own
- Weitzman’s arithmetic: an H100 costs about $30,000 to buy; renting one runs $5/hour for a GCP spot instance or roughly $3.50/hour on Azure or AWS. Annualized, that is $35,000–50,000 — about 1.5× the purchase price per year — against hardware warrantied for three years that he imagines, with uncertainty, may keep working for 10.
- The cultural origin predates the math: renting made Speechify engineers “parsimonious,” afraid of costing the company money. The analogy he and his brother coined: Michael Jordan should not have to pay $20/hour for court time — “you want a hoop in your house.”
- The harder constraint is physical: large-scale training needs a large memory card co-located with the GPU cluster, which cannot simply be rented from Google, Microsoft or even Baseten without expensive commitments and reduced control. Running open-source coding models on owned hardware costs “a fraction of a fraction of a cent per token” instead of paying Anthropic for branded tokens.
- The claimed payoff: Speechify’s newest Simba 3.2 model is “ranked number one in the world for quality,” above the frontier labs, and is 10× more affordable than products such as ElevenLabs.
2. Stebbings’ depreciation objection and the iPhone-drawer rebuttal
- Stebbings’ pushback: chips depreciate, cycles are accelerating, and buying locks the company into one architecture. Weitzman’s answer is that a GPU is not an iPhone gathering dust in a drawer: “if I own 100,000 GPUs, I’m still gonna use all of them at the same time.” Speechify still runs K80s and A100s for inference, where an older card can deliver what is needed in 100 milliseconds; training gets the newest hardware.
- The capital-allocation frame: a strong long-term bond yields around 5%, while buying a GPU can produce a higher return because renting costs 1.5× as much over a year. Excess capacity could also be rented to other users, although Weitzman does not expect to have much excess.
- Demand forecasting is seasonal and layered: if November is 100% utilization, October is 140% and December 80%. Speechify buys the roughly 20% of normal usage it knows it will need, commits another 25% through long-term hyperscaler contracts, and spot-rents the rest without getting close to overcommitting.
3. NVIDIA just built a floor under used-GPU prices
- The structural news Weitzman flags: earlier this month, NVIDIA made a deal with Blackstone, BlackRock, Apollo and Goldman Sachs to support GPU-collateralized lending. If a borrower such as Google or CoreWeave fails, NVIDIA will buy back the GPU for up to 25% of its value. Weitzman says that creates a liquid secondary market and cheaper financing — “exactly what Elon did in the beginning of SolarCity,” when Morgan Stanley and Merrill Lynch amortized solar panels over 30 years.
- On circular-economy fears: what happened between Oracle and OpenAI about a year earlier “was way too much. Like that was ridiculous.” NVIDIA is different, he argues, because the asset has intrinsic value. Unlike Bitcoin, a GPU is useful wherever it is networked; even gold’s uses are comparatively limited to areas such as medicine and jewelry. The question that matters is “how many teraflops per second can this device do?” — effectively the GPU’s token of value.
4. What nobody tells you about physically buying chips — and the “just rent” fight
- The unglamorous specifics: “Dell is a GPU rack supplier at this point,” not merely a PC company. Weitzman’s supplier had Blackwells available in France, but delivery was late — prompting “Pierre, what the heck? We have a contract.” He pays six-figure premiums to skip queues because a late GPU still leaves him paying for data-center space. A truck can carry a house’s worth, or multiple houses’ worth, of GPU value, so insurance matters.
- Rubin systems are liquid-cooled, but many data centers do not already have approved liquid-cooling installations. Speechify researched buying a liquid-cooling “sidecar,” having the data center install it, and then supplying the rack, networking and energy. The energy constraint is now especially important.
- Stebbings’ forceful objection: Speechify is saving roughly 0.5× per year, not 10×, while taking on insurance, freight, cooling and logistics — “I don’t wanna worry about liquid cooling and insurance for a freight truck” while competing with ElevenLabs.
- Weitzman’s rebuttal: ElevenLabs faces the same operational problem; Piotr bought GPUs early and set them up in his house. Renting a co-located cluster can be “ridiculously expensive,” requires multi-year commitments, and provides less control. Buying Rubin systems can give Speechify a year of access before others waiting for hyperscalers receive them.
- His broader framework is that a strong company needs a team, data, compute and architecture. Each engineer may run 5–18 long-horizon agents, and one of his best engineers is focused on creating synthetic data sets rather than writing models. Without compute or data, the team is constrained.
5. Data marketplaces: great business, but it’s not ARR
- Weitzman’s caution on Mercor-type businesses, alongside Micro1 and Surge AI: “it’s not ARR.” Each transaction is one-time, the buyer does not have to buy again, and early investors were skittish for exactly that reason. The provider must be an “operations monster,” move quickly, prove that the customer’s model improves, and indemnify customers on data provenance; “we’ve seen the lawsuits.”
- Stebbings proposes that the customer base could move from frontier labs, which he says provide roughly 90% of current revenue, to large enterprises needing supplemental data for specialized models. Weitzman agrees that data is essential and says companies such as Mercor, Micro1 and Surge AI can shorten the path to revenue, but he focuses his response on the marketplace’s one-off economics and operational burden.
- Weitzman says frontier labs may pay a fraction of the revenue they expect to earn over the next decade for data today, rather than build and manage the collection operation themselves.
6. “100% on me”: how ElevenLabs leapfrogged Speechify
- Asked whether missing B2B was his fault, Weitzman answers “Yeah, 100% on me” and calls it “the biggest strategic mistake I made in the history of Speechify.” He met Piotrek and Mati in London in 2022, admired them and wanted to use their model, but decided that a text-to-speech API would eventually commoditize onto computers and phones.
- What he missed is now his operating theory: an AI lab’s first product is a wedge. One excellent voice can lead to more voices, emotional prosody, voice cloning, speech-to-text, and duplex models that handle ums, laughter, interruptions and turn-taking. From there come voice-conversation harnesses and applications in sales and customer support.
- ElevenLabs first built an excellent API and creator product, then launched Agents, where the buyer becomes a CTO, CIO, CEO or other executive rather than only a software engineer. Weitzman says he forgot the Silicon Valley thesis of getting users onto a product and then selling them additional innovations.
- The economics shaped Speechify’s delay: B2C customers paid less, forcing Speechify to keep costs below $10 per million characters. ElevenLabs charges $100 per million characters; the OpenAI model cited on benchmarks costs $196. Speechify’s newly launched Simba 3.2 API costs $10 per million characters.
7. The B2B debate: Postmates effect vs. “be in the race”
- Stebbings’ case against Speechify’s B2B move is that it now competes with ElevenLabs, which he calls an “unstoppable machine” with government buy-in across major Western democracies, and with Sierra, backed by Bret Taylor, Sequoia and Greenoaks. He frames being third as “the Postmates effect” and argues Speechify could instead remain the dominant consumer brand. He also says value accrues to the top player; Weitzman agrees that power laws are real.
- Weitzman’s counter-precedents: Anthropic was second to OpenAI for a long time and is no longer second; Facebook was second to Friendster and MySpace. The market is oligopolistic rather than monopolistic, and OpenAI — which many would have expected to win voice AI three years earlier — fumbled that niche. Stebbings attributes that to poor management; Weitzman says every company can fumble niches, including OpenAI in coding.
- Speechify’s consumer base is substantial: Weitzman claims 98% of App Store text-to-speech installs, more than 770 billion words served — around 6,000 years of listening — and 60 million users. His plan is to offer products essentially for free, embed them in the user’s stack, and keep innovating. “The best way to lose is not to be in the race. Be in the race.”
8. The talent war: brutal at growth, “the easiest time ever” at seed
- Stebbings’ provocation is that Anthropic and OpenAI offer unusually large, potentially liquid future payouts, drawing people such as the Monzo founder and Matt Clifford. He says there are around 30 heavily funded companies competing for roughly 1,000 highly experienced AI and systems people, with some packages reaching $50 million or more and others at $15 million.
- Weitzman initially distinguishes those companies from ordinary seed startups: a company that has raised $150 million at a $500 million–$2 billion valuation can reasonably offer a $15 million package, but a normal seed founder cannot. He therefore says growth-stage hiring is harder, while true seed companies face “the easiest time ever” because raw technical aptitude can be developed rapidly.
- Speechify now hires math Olympiads, LeetCoders, Kaggle winners and physics or math graduates who may not have coded before. Weitzman says the company can teach them the rest in six months, following a playbook he associates with Duolingo: “hiring for slope more than intercept.”
- Anthropic attracts senior talent because it has built, in Weitzman’s view, “the best, most beloved product for engineers in the history of the world.” It hires many CTOs of public companies and successful startups, more CTOs than CEOs. When Speechify had 21 people, 18 had previously been a CEO, CTO or VP of engineering.
- His practical hiring advice is to use functional interviews, give candidates a large codebase to understand and change, check what they broke, and require competent agent orchestration. On whether people now prefer a certain $10 million from Anthropic to a possible $60 million at a startup, Weitzman says the certainty-risk equation remains individual rather than having fundamentally changed.
9. No credit until it’s in production
- The signature culture analogy: “Imagine you’re in the milk delivery business, and you make me a beautiful bottle of milk, and you leave it down the road. The milk will spoil.” Credit accrues only when features reach real users without bugs. “We are not in the theory space. We are an applied AI company. That’s why we win.”
- His proof point is a 19-year-old engineer who, while waiting for three training runs to finish, implemented all 14 of Weitzman’s recorded product notes. The result was demonstrated the next morning and solved the problems Weitzman had identified.
- The stack is Claude Code first, Cursor second and some Codex. Linear tickets can be handed to agents; a good engineer is now “an exceptional QA” who tests the feature, identifies edge cases, prompts fixes, optimizes the result and makes roughly 10 product and architecture decisions a day.
- Agents require close supervision: Weitzman says his brother Tyler once set an alarm for 3 a.m. to check what an agent was doing and describes babysitting an agent roughly every three hours. He also warns that this level of engagement can create AI fatigue and burnout.
- Token discipline does not use leaderboards, which Stebbings calls a bad incentive. Speechify judges demos and production outcomes, and Weitzman says people can be let go for wasting tokens on trivial work. By contrast, he cites an Anthropic example involving Fable 1 in which a long-horizon run spent $12,500 over two weeks to produce a better model. That is exactly the right use, he says: define the target and measurement, then iterate against it.
10. “Only losers compete”: Wispr Flow/WhisperFlow, Siri and the wedge relearned
- Stebbings describes a deleted post about Wispr Flow, later called WhisperFlow in the discussion, becoming worse and prompting 500 alternatives — “talk about the commoditization of a market.” Weitzman says the product originally used a harness combining several components, probably DeepL or Deepgram under the hood, and may have worsened after switching to a cheaper in-house model.
- He also argues that public announcing can attract competition. WhisperFlow has a bigger technology-world brand because it announces more, while Speechify intentionally does not; Weitzman claims Speechify has far more users and no comparable competitor at its scale. He cites Peter Thiel: “Only losers compete. Try not to compete.”
- Weitzman made a similar mistake with speech-to-text. He built his own experience seven years ago but assumed Apple would add the feature natively, just as Apple’s announced ChatGPT–Siri partnership produced no visible result. Speechify has now launched products competing with Siri, Wispr Flow/WhisperFlow and ElevenLabs.
- On Stebbings’ skepticism about customer support — including his claim that 18 companies raised more than $100 million in 18 months while sophisticated buyers build their own systems — Weitzman says Speechify’s B2B core is the API, not a generic customer-support product. He claims an API advantage in quality, speed and price, while forward-deployed engineers work with customers to build agents and discover the next product.
- Sierra and ElevenLabs: Weitzman says both will be massive and says he would not try to fight Bret Taylor, whose résumé includes Google Maps, Facebook CTO, Salesforce co-CEO and the OpenAI board. Stebbings distinguishes Taylor’s broader, tool-calling and Salesforce-like strategy from ElevenLabs’ voice-centric strategy; Weitzman agrees the companies are playing different games but says both are pursuing the broader AI-agent opportunity.
11. Voice-first computing, and buying Meta over xAI
- The five-year contrarian call is that the human-computer interface becomes primarily voice. Google and ChatGPT succeeded partly through simple interfaces — “a text box and a button” — and conversation is simpler still. Weitzman says current ChatGPT voice is too slow, uses a weaker model and escalates poorly to the higher-quality model. He thinks Meta has the right idea with voice-enabled phones and wearables.
- Forced to choose between xAI and Meta, Weitzman chooses Meta: “Elon’s distracted.” He then evaluates Elon through SpaceX and Tesla, setting aside space data centers temporarily. He says Tesla’s energy storage and chip-manufacturing work matter as memory and energy constrain data centers, while Elon’s space-data-center idea was “a great rabbit out of the hat” for pitching investors. Stebbings says it “ruins all estimates.”
- Meta, Weitzman argues, has more data than anyone, though GDPR and other U.S. laws constrain its ability to train on that data. Meta trades around a 32 P/E, while Tesla trades in the multiple hundreds and SpaceX is valued “insane[ly].” Zuck has roughly 20 additional years of potential leadership compared with Elon.
- Stebbings counters that removing Zuck could produce short-term stock-price appreciation by replacing the “CapEx, CapEx, CapEx” focus with an ads-business executive, whereas removing Elon would destroy much more value. Weitzman first says Meta would be dead without Zuck, then emphasizes that Meta lacks the Alex Karp or Elon narrative-premium effect. Its underlying technology and business value may therefore be much larger than its market valuation, unless Elon’s larger vision succeeds.
12. GPUs against orphan disease: the personal stakes
- Weitzman’s most exciting frontier is AI in pharmacology and biology. For a family member with severe autoimmune neuroinflammation, he took weekly blood samples for 15 weeks, ran genome sequencing, proteomics and RNA analysis, and compared those results with six years of self-reported quality-of-life and mood data on a GPU cluster. He says he found things no doctor had been able to tell him.
- He is buying a roughly $5,000 pocket-sized sequencing device and plans to organize meetups with people who have the same rare disease, including members of its Facebook group. He wants to sequence their genomes and compare them on a large GPU cluster to find an epigenetic common thread: “I know I’m gonna solve this disease.”
- He then describes using conclusions from that work with AlphaFold from Isomorphic Labs, designing molecules that bind to relevant proteins, using CRISPR, and ordering RNA or DNA sequences from Twist. He says the work can be simulated through an SSH connection to his GPU cluster in Scottsdale, Arizona.
- He also says GPUs helped him identify the location of his father’s prostate-cancer lesion. The through-line back to Speechify is personal: his father read Harry Potter to him when dyslexia made reading difficult; after moving to the United States at 13, he listened to the audiobooks 22 times in a row; and he built a text-to-speech tool that helped him graduate from Brown. “Technology solved my dyslexia, and it solved my ADHD, and it’s gonna solve my brother’s disease.”