Turing CEO Jonathan Siddharth: Who Wins in Data Labelling & Why 99% of Knowledge Work Will Disappear
Turing CEO Jonathan Siddharth: Who Wins in Data Labelling & Why 99% of Knowledge Work Will Disappear
Summary
- Jonathan Siddharth’s core claim: “the era of data labeling companies is over and it’s now the era of research accelerators.” Three shifts drove it — data went simple to complex (“write a Python program to sort some numbers” became “write a B2B marketplace app” across Kotlin, Swift and Next.js), training went from passing tests to doing real work, and chatbots became agents needing RL environments; chatbots used SFT/RLHF, while agents add reinforcement learning. Turing trains “superintelligence” for seven of the eight frontier labs and now builds “a mini world model for business” for every workflow in every role, function and industry — “that’s like $30 trillion of knowledge work.”
- On the bubble question, a flat no: “I don’t see an AI bubble… GPT-5 is like f*ing awesome. We’ve just gotten used to magic.” The tradeable idea is the model capability overhang — “the models are capable of X but what we are getting out of the models is X minus delta” — meaning value gets unlocked by scaffolding, evals and deployment, not just new pretraining runs. MIT’s 95%-of-pilots-fail stat is growing pains, not a wall.
- The revenue-vs-GMV debate gets a careful answer: Turing describes its revenue as “gap” or traditional revenue numbers, and “these are not SaaS ARR numbers… this is a different beast” — recurring lab projects sustained only by performance. Concentration is Nvidia-like by design: labs spend with “a small handful” of trusted, firewalled partners for resilience, and Harry cites Nvidia taking 39% of revenue from two clients at a $5T market cap as the comp. Scale’s acquisition “flooded” Turing with demand.
- Harry’s strongest pushback: enterprises “are so far off adopting Slack and Notion, let alone building custom models” — automation in 20 years, not 10. Siddharth’s split: back office slow, front office fast, because “it’s a lot easier to convince people to use a piece of technology to make more money than to save money” — and OpenAI’s GDPval showed the best models (likely Claude 4 Opus first, GPT-5 also quite good) producing work “indistinguishable from a human expert” about 50% of the time on real single-step tasks.
- “SaaS as we know it, I think is over. It’s completely over” — three kill vectors: companies build custom apps themselves, foundation models “sonic boom” the app layer, and GUI use goes away as ambient AI may use MCP and tool calls. Harry dissents hard — companies run 80-100 SaaS products, the long tail “can barely use Wix and Squarespace,” and verticalization defends (“Sam is not going there” on patent software). Siddharth’s retort: audit your portfolio — today’s startups likely use fewer SaaS apps and fewer people.
- The AGI-pilled game theory behind circular deals: whoever wins superintelligence “will probably win search… consumer devices… operating systems… cloud… social networking,” so Zuck spending $100B (12-18 months of free cash flow) against a $2-3T market-cap downside is rational. Both agree: “You have to play.”
- Where he’d invest in his own space: “probably in robotics or embodied AI” — vertical data acquisition is innings one but no longer white space, while robotics data is “wide open.” He believes in slow takeoff, not rapid — good for the world because unlike self-driving, AGI unlocks “incremental value for every percentage improvement” — and sovereign models are coming: Harry sees no way German healthcare runs on American models, and Siddharth agrees governments will need their own nationals generating training data.
Deep dive
1. Data labeling is dead — “it’s now the era of research accelerators”
- Siddharth’s definitional move up front: Turing is not a talent marketplace — “we’re training superintelligence.” Superintelligence needs research (labs do in-house), compute (“we have Jensen to thank”), and data — “Turing powers the data pillar” for seven of the eight frontier labs.
- The first shift is simple to complex data: a few years ago coding data was “write a Python program to sort some numbers”; today it’s “write a B2B marketplace app that connects doctors with patients” — in Kotlin for Android, Swift for iOS, Next.js for web. “It’s no longer the kind of data that low-skilled, medium-skilled contractors can generate. You need expert humans in every domain.”
- Shift two: from passing tests to doing work — “it’s less about having AI pass the bar. It’s more about can AI do the job of a lawyer” — a privacy lawyer, a compliance lawyer, a paralegal. Shift three: chatbots to agents, which changes the data entirely — chatbots used SFT plus RLHF; agents add reinforcement learning in RL environments. Hence the thesis quote: “the labs want to work with a proactive partner that can think about what types of data are likely to be helpful.”
2. RL environments are mini world models over $30 trillion of work
- The SDR example, as told: clone LinkedIn, Salesforce and ZoomInfo with synthetic databases, prompt the agent to “prepare for a call with this person… after the call update Salesforce,” and let a verifier check completion while the agent tries trajectories and tool calls. Curriculum design is the craft — too easy or too hard and “the model doesn’t learn much.” “It’s very similar to the technique that AlphaZero used in mastering Go.”
- The scale claim is a four-dimensional matrix — every industry × every function × every role × every workflow: “we are creating RL environments for every workflow for every role in every function in every industry. That’s like $30 trillion of knowledge work.”
- A competitor’s board member told Harry “the big thing we all got wrong was we are so in innings one of the acquisition of verticalized data.” Siddharth: “Absolutely. It’s innings one” — and the whole RL-environment regime is only ~12 months old, sparked when “o1 dropped in December, DeepSeek launched in Jan.” One year later “it could be something totally different,” which is why he says the market rewards research DNA.
3. Enterprise reality: small on-prem models and the “first mile schlep”
- The insurance-underwriting specimen: unstructured medical data in, risk tier out. You don’t need a trillion-parameter world model — a half-billion to 10-billion-parameter model, on-prem, fine-tuned on a decade of proprietary underwriting judgments, is faster, more accurate, and keeps data away from frontier labs. This is the forward-deployed business with Disney, Pepsi, BlackRock, Fiserv, and Johnson & Johnson — smaller than the lab business but “growing pretty fast,” and priced by time for now: “I don’t think that’s the right way to do it,” value-based pricing later.
- First mile schlep in the wild: “our data is a mess. It’s in silos… some of the data is in a file that Bob has and Bob doesn’t work here anymore.” Then evals, a cursor-like interface designed for partial autonomy, and training humans on new workflows.
- Deployments run as a tandem system — human and AI do the same job while a manager compares output; human errors train the human, while agent errors become data for fine-tuning the next iteration. Harry: “If the agent is right and the human is wrong, why don’t you just fire the human?” Siddharth: you track precision and recall over time — “you wouldn’t fire them over a single mistake.” Harry: “Bit harsh.”
4. The adoption fight: incumbent decay vs the front-office wedge
- Harry’s pushback in full — enterprises are “so far off adopting Slack and Notion, let alone building custom models… maybe in 20 years, but not in a 10-year time frame.” Siddharth’s counter: if a competitor operates with “100th the headcount” while pricing insurance better, “they’ll get their lunch eaten.” Siddharth extends it into a “10 to 20 year decline of incumbents… transfer of value from old incumbent to startup — hence why we invest.”
- Siddharth’s hypothesis: back-office automation will be slow; the front office moves first, especially financial services — “it’s a lot easier to convince people to use a piece of technology to make more money than to save money.” He cites Mark Chen at OpenAI: financial services is the bleeding edge of the S&P 500 — yet still “about two years behind the state-of-the-art.”
- Rory O’Driscoll’s test, via Harry: AI value hinges on whether budget transfers from human labor to AI technology. Siddharth says the transfer is already high in customer support, copywriting and SEO — “low-risk-to-fail areas.”
- His evidence for the long arc is GDPval, OpenAI’s study that he recalled as covering 9 verticals and 44 occupations doing real deliverables: “about 50% of the time the best models were producing work that was indistinguishable from a human expert” — the number one model was likely Claude 4 Opus (“kudos to OpenAI” for flagging a rival), while GPT-5 was also quite good — though on single-step tasks. “We are well on our way to AI eventually automating all types of knowledge work.”
5. Revenue is “a different beast” — GAAP, firewalls, Nvidia-grade concentration
- On the GMV-vs-revenue controversy he declines to name names but draws his own line: “we think about revenues… in terms of gap revenues,” meaning traditional revenue numbers. “These are not SaaS ARR numbers… this is a different beast” — recurring lab projects that start and end, with “lots and lots of demand” but only “as long as you’re doing a good job.”
- Trust is the operating constraint: projects are firewalled between labs and even between teams within a lab — his analogy is Foxconn’s floors, iPhone made on one, Pixel on another. Labs deliberately keep “a small handful” of partners for resilience: “we know what happened when the Scale investment happened.”
- Scale’s acquisition: “We just got flooded with a lot of demand” — and Turing “amped up pretty significantly in multimodality,” Scale’s strength from its autonomous-labeling roots. His most-respected competitor is Alex Wang, “prescient in seeing the importance of data.”
- On seven-customer concentration risk: Siddharth says, “I think we are in the same boat as Nvidia” — while Harry supplies the comparison that Nvidia has 39% of revenue from two clients and ~50% from four, “for a $5 trillion company.” Stargate alone is “a $100 billion a year investment on compute,” and the next demand leg is sovereign: Harry — “I do not think there’s any way you’ll have the German healthcare system working with American model providers” — and Siddharth agrees, with governments wanting their own nationals generating SFT and RL data.
6. “I don’t see an AI bubble” — the overhang and the forced $100B bet
- The categorical position: “These models are incredibly powerful today… GPT-5 is like f*ing awesome. I think we’ve just gotten used to magic” — and “they’re the worst they’ll ever be.”
- His central mechanism is the model capability overhang: “the models are capable of X but what we are getting out of the models is X minus delta” — closed by agentic scaffolds, context engineering and tool access. Live demo: Harry’s 12 hours every weekend picking 15-20 clips per show is exactly what a fine-tuned scaffold could do. Harry: “If you could f*ing make it work, dude, I’d pay you a lot of money.”
- The MIT 95%-of-pilots-fail stat is growing pains, not falsification: unstructured data, no scaffold, no evals, no partial-autonomy workflow — he cites Karpathy on why Cursor works. Some roles skip straight to full autonomy (customer support — internet tokens suffice); others don’t: the way one firm does financing might look different from another’s, so fine-tuning is required.
- On circular deals signaling a bubble: if you’re in the AGI camp — and he is — the winner of superintelligence “will probably win search… consumer devices… operating systems… cloud… social networking.” Harry runs Zuck’s math: spend $100B (12-18 months of free cash flow) and fail alongside everyone, fine; don’t spend and someone else wins, lose $2-3T of market cap. Both, in unison: “You have to play.”
7. “SaaS as we know it, I think is over” — met by Harry’s best pushback
- Three kill vectors: apps are now trivially easy to build on LLMs so companies build custom software themselves; if an agentic model is sufficiently integrated into the org’s database, you might get “sonic boomed by the foundation model companies” and “you don’t need anything else in the middle”; and the GUI may go away — “the GUI was designed for a world where humans were using a keyboard and a mouse… humans can do better things with their time than click around.” Ambient AI may use MCP and tool calls instead.
- Harry dissents at length: the average company runs 80-100 SaaS products — nobody will build and maintain that internally; the long tail of “every plumbing provider, law firm, accounting firm… can barely use Wix and Squarespace”; and verticalization defends — on AI for patent creation, “Sam is not going there.” Siddharth’s empirical retort: tally SaaS usage across Harry’s own portfolio pre- and post-ChatGPT — “my hypothesis is that today’s companies use fewer SaaS apps and have fewer people.”
- One moat he sees is data-driven feedback loops. PageRank’s recipe was known across Google, Yahoo and Microsoft; Google won because user preference produced representative queries and clickstream — “a high-quality gradient for which direction to step in.” Enterprise is still “wide open”: deploy first, discover where models break, generate data to plug the gap.
8. After automation: intelligence as an API, more engineers, digital surrogates
- Three consequences of full knowledge-work automation: 100x leverage (“Elon maybe runs 600 companies”); an entrepreneurship boom — the therapist founder recruits “a marketing GPT, a software engineer GPT, a PM GPT” instead of raising a few hundred K; and “a million flowers will bloom” far beyond London and Palo Alto.
- Harry’s darker read: 6.5 million people in the UK’s working population don’t work, and “we grossly overestimate the intelligence of the general population” — won’t this widen the chasm? Siddharth flips it: superintelligence is “intelligence as an API”; for $20 a month, if available, it could provide access to expert intelligence versus the expensive human expert who is the real gap-widener. And no beach: “we are tool builders… we’ll solve problems at higher and higher levels of abstraction” — cure diseases, reverse aging, “go to the stars.”
- More software engineers in 10 years, not fewer — the definition expands to anyone who ships software that solves a real problem, like his Stanford oncologist building a home-diagnosis app. Discovery amid infinite software resolves via agents talking to agents, Her-style: Harry might be having “a million conversations with entrepreneurs” at once.
- Hardware follows: an always-on wearable with cameras and an earpiece that whispers “Harry seemed less interested… but when we were talking about AR he perked up.” The phone survives diminished: “the phone app is like the least interesting part of the phone.”
9. The ten-year map — and what he changed his mind on
- Market structure: “a few winners,” not a monopoly — labs want resilience and price competition among partners, and the market rewards research depth because paradigms flip yearly. Where he’d deploy capital in his own space: “probably in robotics or embodied AI” — vertical data is innings-one but no longer white space (Turing is scaling it), while robotics data is “wide open,” home robots needing different data than factory ones.
- Quickfire convictions kept as hedged: slow takeoff, not rapid — and that’s good, because AGI is not self-driving where the unsolved last 1% kills usefulness; “there’s incremental value that’s unlocked for every percentage improvement.” On China: the frontier circles he works in “don’t underestimate” it — DeepSeek, Kimi K2, Qwen are “state-of-the-art.” Frontier models carry “some value in keeping some of the technology closed”; enterprises will mix open and closed in the 0.5B-10B small-model regime.
- The changed mind: he used to believe in hiring “a strong exec team” and getting out of the way; now it’s Elon-style ground truth — walking the factory floor asking “why this door in the Model 3 has three bolts instead of two.” The confession attached: “in the early days of starting Turing… I may have had a subconscious desire to be liked.”
- Most unpopular decision: switching from the distributed team to hub-and-spoke — SF, Palo Alto, a London office coming; “some of them left.” Harry shares that his mother has MS and thinks there will be breakthroughs in MS drug discovery. Siddharth says what excites him most is “automating AI research itself” into a self-improvement loop. The closing image is Iron Man’s agentic drone suits: “Today Harry might have a hundred ideas, but Harry’s able to do maybe two of them really well. I like a future where Harry can do the remaining 98.”