Pioneers Insight Method Research Author
Factory's Reyes: Anthropic's $2T coding bet; 80–90% of neo-labs die
Back to Episodes

Factory's Reyes: Anthropic's $2T coding bet; 80–90% of neo-labs die

Summary

  • Eno Reyes’ core pricing thesis is that AI should be priced by outcomes, not tokens — and on that math “the smartest model is actually the cheapest.” A frontier model that nails a code review in 1,000 tokens beats a cheap model burning 50 million to get there, which breaks the “a token is a token” framing and points to rapid “speciation” of models: open commodity models for everything, plus post-trained internal specialists enterprises keep entirely to themselves.
  • He thinks “the TAM of frontier models is frankly over-weighted right now,” and Harry frames Anthropic at $2 trillion as effectively a $2T price on Claude Code in “one of the most competitive application markets in one of the most finicky segments” — dev tools. Baked into $2–4T valuations is the assumption labs can “2X the price of those tokens and people will buy them”; the escape hatches are regulatory capture or figuring out how to build applications and outcomes that match the cost-quality frontier, likely by opening up to more models, which OpenAI is quietly doing.
  • “It could be 80 to 90% of neo labs die in the next 18 months” — though “die” often means good acquisition outcomes rather than doom. His three-question durability test: is the workflow durable, does it survive better frontier models, and does it survive an entirely new way of working? Legal passes all three; “general computer use” and Excel/Jira-adjacent knowledge work fails.
  • “Calling open source models Chinese models is a psyop by the frontier labs to basically trick people into thinking that they’re scary and otherize them.” His prediction: “In three years, 99% of workflows are gonna be done on open models. But 1% of those tasks is probably gonna be 30, 40% of the economic value” — frontier use cases such as bio research, defense and advanced AI development are “incredibly niche,” so cost dominates for the Global 2000.
  • Routing technology is commoditized; Stripe’s $8B OpenRouter buy was a bet on capital-allocation information, not tech — “the technology’s just no longer the moat.” The real leverage sits in the harness, “effectively the new sort of application”: context windows were solved by compaction inside the agent, closed-loop continual learning “has not been developed. It doesn’t exist,” and learning accrues at the harness layer.
  • The defining question of the next five years: “Who is the sovereign of your intelligence? Is it you, or is it some other company?” Two of the largest model companies “have explicitly said, ‘We’re going to go after every single one of these industries,’” which drives on-prem demand (Factory Private) and makes Cursor’s SpaceX tie-up a liability — hard to stay model-independent when “they’re gonna wanna push Grok.”
  • Marry Microsoft, shag NVIDIA, kill Meta — and NVIDIA at $10T in three years is “likely yes” if SpaceX gets to be worth $2–3T. Microsoft is “one of the best-positioned hyperscalers” thanks to model independence and infrastructure; the debt cycle is survivable for hyperscalers but “totally existential” for OpenAI/Anthropic, who “need to become the single greatest free cash flowing businesses in the history of technology in order for them to just live.”
  • On talent: “We expect 100% of our future hires to come through acquiring companies,” Ivy League pedigree “is barely a signal for competence,” and performative 9-9-6 culture is a red flag “almost always correlated with making up for some other detractor.” Token spend should map to projects, not people — Factory put “almost seven figures of credits in one day” against one benchmark (Program Bench), and sees such budgets reaching “eight and nine figures easily.”

Deep dive

1. Price the outcome, not the token — the smartest model can be the cheapest

  • Eno’s opening frame: “How much does a code review cost is far more interesting than how much do the tokens inside of that code review” cost. A sophisticated model that gets the outcome right in ~1,000 tokens beats a cheap model that grinds through 50 million — “for many of the most demanding tasks, I see a world where the smartest model is actually the cheapest.”
  • Harry’s challenge — doesn’t this contradict “a token is a token”? Eno’s answer is model speciation: open models dominate commodity task execution, while businesses with a few high-volume specialized tasks — where commodity isn’t good enough and frontier is too expensive — post-train a middle model and “keep it entirely internal.” Not millions of models, “but it’ll definitely be quite a lot.”

2. Post-training gets democratized like software did

  • Harry’s objection: company structures today simply aren’t equipped for post-training and implementation. Eno’s rebuttal: the recipe currently “lives primarily in the heads of a specialized few” — exactly where software development sat 20 years ago. Soon an enterprise will “open up a platform, click a couple buttons, describe the task you care about, point it towards those workflows… and out comes a model.”
  • The implicit shot at the labs: “a couple of companies claim that recursive self-improvement and model training will be only their domain, and I think in reality, many businesses will have access to that technology via software services that other companies sell.”

3. The frontier of AI is building verification where none exists

  • Verifiability is “the single most important property of success with current AI systems.” In fuzzy domains like healthcare and legal, eval creators bring in experts to construct new forms of judgment — and the next step: “the frontier right now of AI is AI systems that can build verification where there is none,” letting them advance into tasks humans consider too hard for AI.
  • His management analogy: a novice manager at a law firm judges hires by gut; at scale you must write down what good looks like — but that changes behavior. “If you say good looks like A, B, and C, you’re gonna get a lot of A, B and C,” whether or not the incentives were right.
  • Harry’s echo — “show me the incentive and I’ll show you the outcome”: set a VC’s goal at three deals per year and you get three deals, not a great one.

4. Harry’s Mercor regret and the order-of-magnitude error

  • Harry’s confession: he invested at ~$2–3B, skipped the ~$20B round because “how much bigger can it be? Maybe 100 billion, but that’s a 5X” — and now thinks he’s “completely fucking wrong,” seeing a path to $200–300B on future data requirements.
  • Eno agrees the market misreads scale: people calling $20/30/50B valuations ludicrous are “underestimating by an order of magnitude how massive a transformation this is gonna be.” But the winners won’t look like moat businesses — they’re “collections of people that understand what the future looks like a little bit more clear-eyed than the other people.”

5. Frontier TAM is over-weighted — and Harry frames $2T on Anthropic as a bet on Claude Code

  • “The TAM of frontier models is frankly over-weighted right now.” The world assumes 1–3 companies dominate the intelligence era; the real question is defending margins amid so many options, when $2–4T valuations assume you can “2X the price of those tokens and people will buy them.”
  • Harry’s pushback: Anthropic just posted its first profitable quarter with ripping margins. Eno: the driver is applications — model margins are “definitely worse than the applications.” The lab strategies split: dominate the inference platform or move up-stack; “Anthropic seems to be following the application path while OpenAI seems to be dipping its toes in both.”
  • Model-lock is “bad incentive alignment”: a locked provider gives you their best model, not the best model. Harry presses — isn’t $2T “a $2 trillion price on Claude Code,” which is “not that difficult to switch off of”? Eno concedes that’s exactly the risk: “a $2 trillion bet on one of the most competitive application markets in one of the most finicky segments… dev tools.”
  • The two escape hatches are regulatory capture or figuring out how to build applications and outcomes that match the cost-quality frontier, which Eno says means opening up to more models. OpenAI is grappling with this: “they’re not making it official, but they’re clearly supporting an open model ecosystem.”

6. The worst marketing job in contemporary capitalism — and revised predictions

  • On Dario’s messaging: AI marketing “did the opposite of what you want. Scare every single person, tell them it’s very unreliable, basically threaten their wellbeing and livelihood with the technology while you roll it out at scale.” The threats are real, but “the moment you start talking about the singularity and AGI and create this godlike mythology out of AI, you’re gonna scare a lot of people.”
  • Eno gives Sam Altman credit for revising a prediction: “I totally underestimated the momentum of the economy… so I predicted this future that actually has not come true.”
  • Eno’s own corrected prior: “building a business is much more reactive than planning” — the best decisions were split-second reactions, and the future is defined “in real time in group chats.” The upside: “you’re basically a couple of decisions away from even greater outcomes.”

7. Margins, subsidies, and Factory’s refusal to play the flood-the-market game

  • “Not all businesses in AI have bad margin profiles. I mean, I know we’ve got good margins.” Factory doesn’t subsidize consumers in dev tools — costly in mind share — because “there are two players with effectively infinite money who are trying to flood the market,” and “you don’t keep them after you pull back the subsidies, as we’ve learned.”
  • The long game: the most cost-effective solution in 1–3 years is open models running locally, so Factory optimizes for local, open, and on-prem now — “eventually the self-service will come to us… because we have the best product in market.”
  • Factory still has a steady stream of tens of thousands of daily self-service users. Eno says fewer than 250,000 may be enough to create the feedback loop needed for product improvement, though the grassroots community, media and storytelling that come with millions of users are harder to reproduce.
  • Investor advice: subsidy-for-lock-in is a classic strategy, but “if there’s no path to increasing the margin profile, that is very risky” — raise price without raising outcomes and “people will churn and move to another thing.” Also worth keeping: “if you have a product that requires 100 FDEs to get it deployed, you just have a bad product.”

8. Routing is commoditized; the harness is the new application

  • Harry’s puzzle: OpenRouter sells to Stripe for $8B, yet everyone — Ramp, Merged.dev, Requesty — does routing. Eno: the tech isn’t differentiated; Stripe bought a bet on capital allocation. Energy becomes intelligence becomes dollars, Stripe already controls the money flow, and OpenRouter shows “where these models are going.” “You wouldn’t pay $8 billion for the same company that had no users with better technology… the technology’s just no longer the moat.” Still: “$8 billion is quite steep… eight and 10 billion’s the new one billion.”
  • Gateway routing yields 10–20% cost savings, but agentic workflows need stateful, in-task intelligence allocation — like context windows, solved not at the model endpoint but inside the agent “with something called compaction.” “People really want the problems to be solved somewhere else, like in the model or in the gateway, but more and more we see it’s the harness that solves these problems” — “the harness is effectively the new sort of application.”
  • On continual learning: the closed-loop LLM version “has not been developed. It doesn’t exist.” Instead, even model providers acknowledge learning happens at the harness layer — which businesses will insist on owning.

9. Sovereign intelligence, on-prem, and Cursor’s SpaceX problem

  • The five-year question: “Who is the sovereign of your intelligence? Is it you, or is it some other company?” His example: a law firm outsourcing every case to ten vendors finds, five years later, “those other companies can just turn around and screw you over.” And it’s not paranoia — “two of the largest companies that provide models today have explicitly said, ‘We’re going to go after every single one of these industries.’” Palantir and Satya have been loud about owning your intelligence.
  • On-prem is about the idea of control, not the technology: Factory Private is one of their most popular offerings, yet many buyers choose SaaS because they have the peace of mind of knowing they can switch if needed.
  • On Cursor/SpaceX: great outcome for the team, but staying model-independent while attached to a model lab will be a very hard story — “they’re going to wanna push Grok” — plus enterprise trust concerns. Enterprises will take “a second look” at ceding their software development lifecycle to a model-locked provider.

10. 80–90% of neo-labs die in 18 months — and capability progression lags by sector

  • Harsher than prior 20VC guests: “I think it could be 80 to 90% of neo labs die in the next 18 months” — though “die” will often mean “incredible outcomes” via acquisition; “these businesses may not make sense as independent businesses.” The durability test: is it attached to a durable workflow, does the workflow survive better frontier models, and does it survive “an entirely new way of thinking”? Legal passes all three; “general computer use” and Excel/Jira intermediate knowledge work fails — “we may not use a lot of tools like that in five to 10 years.”
  • On Harry’s clipping-this-podcast example, sectoral lag is real and multi-year: media businesses haven’t extracted the tacit knowledge — knowing where the hook is “goes to taste,” an intuition people “couldn’t even describe” — and “the moment businesses capitalize on that delta, the progression will happen extremely quickly.”

11. “Chinese models” is a psyop — and 99% of workflows go open

  • “Calling open source models Chinese models is a psyop by the frontier labs to basically trick people into thinking that they’re scary and otherize them.” Ask the same three questions of every model: what’s censored, does it solve your problems, can you switch if it disappears. Chinese frontier models “have demonstrated no examples” of security backdoors versus American models — just creator bias. His pointed example: write a 10-K describing a strategy involving recursive self-improvement and “it will block you. The answer is Anthropic.” Caveat kept as stated: US national security work “definitely should not be using Chinese models.”
  • Eno expects model creation to continue for quite a while and perhaps accelerate: building models gets easier, while sovereign intelligence produces models with different opinions and perspectives. Model routers also function as an information stream and free advertising whenever a new model drops.
  • The headline call: “In three years, 99% of workflows are gonna be done on open models. But 1% of those tasks is probably gonna be 30, 40% of the economic value of the future of intelligence.” The frontier use cases Sam and Dario tout — frontier bio, LLM development, defense — are “incredibly niche”; ask anyone in the Global 2000 what they’re doing today and “it’s not something that needs the true frontier of intelligence 99% of the time.” So cost dominates — and the labs releasing open models they then inference “could actually be a great business for them.”

12. Microsoft over Meta, existential debt, and the Yahoo era

  • “Microsoft might be one of the best-positioned hyperscalers, honestly, with respect to AI because of this independence” — Satya “played a masterful game” (Kevin Scott sourcing the OpenAI deal), captured the upside, and now sells Azure as home for Anthropic, OpenAI, and open models alike. Zuck’s open-model push is “right for humanity,” but Meta must power its own operations with those models to justify the spend. Forced choice: “Microsoft for sure” — the infrastructure means “no matter what model runs on top of that, Microsoft is gonna win.”
  • On the debt cycle: dangerous without free cash flow. Hyperscalers survive a hit to projected AI cash flows; for OpenAI/Anthropic it’s “totally existential… they need to become the single greatest free cash flowing businesses in the history of technology in order for them to just live.” Chips (the “Jalapeño” chip, Anthropic’s reported effort) make sense — verticalization is the strongest way to free them from massive debt and burden — but the explicit goal is replacing current vendors: “it’s coopetition with everyone.”
  • On froth: no 2008-style crash — but “we are probably in like the Yahoo era,” with OpenAI and Anthropic closer to Netscape analogies; the Jobs lesson Factory prefers is “not being first but being the best.”
  • On Airtable at ~$2.5B (down from $11B): “contemporary SaaS businesses are more like movie studios” — one blockbuster isn’t enough, “if you rest on that, then yeah, Bending Spoons will come and eat you.” Expect a big M&A wave of good-but-not-Stripe businesses.

13. Talent: acquire founders, kill performative culture, allocate tokens to projects

  • “We expect 100% of our future hires to come through acquiring companies.” After Harry’s pushback on the phrase “mission-aligned,” Eno defines it: mission alignment “isn’t a property that you can suggest or say. It’s actually extremely evident in the work” — people already building software-development harnesses with tens of thousands of daily users. On pedigree, in response to Harry’s mention of Cognition’s chess champion and math prodigy: “I went to an Ivy League school. I learned firsthand that that is barely a signal for competence. There are plenty of idiots who went to Ivy League schools” — hire for operating outside what the system calls the rules.
  • Biggest founder hiring mistake: “performative work culture” — 9-9-6 signaling is “almost always correlated with making up for some other detractor.” Harry clarifies his own 9-9-6 reputation: it means jumping on a Sunday client call, not literal hours; Eno agrees but insists “any time you create an incentive to show people that you’re working rather than to actually do work, you’re basically incentivizing the wrong thing.”
  • On token budgets (contra Jason Lampkin’s “$100,000 of tokens to our best engineers”): allocating credits to people “is a very weird way to think about it” — allocate to projects and outcomes. Factory put “almost seven figures of credits in one day” against Program Bench via work being done by one person, and sees such spend “approaching eight and nine figures easily” for some businesses.
  • On pricing heads, after Harry’s Poolside example involving employees moving to NVIDIA at a rumored $12B: “it’s not that any one node is worth $100 million, but when you put all of these nodes together, the graph they make can be worth tens of billions” — for a pre-AI incumbent, the right graph update “can be the difference between a $2 trillion company being a $4 trillion company. So almost anything’s worth it.”

14. Quickfire: threat board, Salesforce buy, and the priestly class of 2 million

  • Shag/marry/kill on Meta, Microsoft, NVIDIA: “marry Microsoft, shag NVIDIA, kill Meta” — NVIDIA is “the current kingmaker,” Meta is “technologically correct” (VR, open models) but has “one cash cow.” NVIDIA at $10T in three years: “if we let SpaceX be worth 2 or 3 trillion, then Nvidia probably is worth 10… the answer is likely yes.” And he’s a buy on Salesforce: “when people say, ‘I hate that software,’ and everyone buys it, that’s probably a pretty good business” — though agile itself may get hit, opening a next-system-of-record threat to both Linear and Atlassian.
  • Threat ranking: Claude Code first (“brought up in every single conversation”), Codex second — increasingly “Codex for Work,” and switching off Claude Code “is actually the greatest news for us because it shows how basically unsticky this is” — then Cognition (basically the only other model-independent enterprise vendor), then Cursor (perceived as an IDE, not an enterprise strategy). The differentiation: almost all four pitch an eventual human-level AI labor replacement — an “Indiana Jones swap” replacing people in the business — while Factory pitches “an entirely new development methodology” humans build alongside AI.
  • On Chamath’s “Silicon Valley is too money-centric”: “That is something that it sounds like Chamath would say because frankly, that’s how he’s made his money.” San Francisco is full of people eating glass on hard problems — and the stories we tell matter because “that technology is built on the stories that we tell.” On Chamath’s software venture, success “depends on how real the software is.”
  • Enterprise sales advice: “stop treating it like persuasion… treat it as a discovery opportunity.” And the closing five-year prediction: it will seem unthinkable “that we let a sort of priestly class of maybe 2 million people decide the fate of all software for all of humanity” — a boat operator in Belize will have fully custom software “that looks better than your HR IT software back at home.”