OpenAI vs Anthropic vs Open-Source | Token Maxing, AI Hangovers & The Coming ROI Reckoning
OpenAI vs Anthropic vs Open-Source | Token Maxing, AI Hangovers & The Coming ROI Reckoning
Summary
- The model war likely has no single winner. Matan’s biggest change of mind in 12 months: he used to think one or two labs would run away with the frontier — now “it’s probably going to be at least four that are going to probably be approximately as good. And that is a win… the bad case for humanity is when there’s one that’s really really good.” Factory’s own bear case is the mirror image: it dies only if one lab pulls decisively ahead — “but then that’s a monopoly for the entire economy to be worried about.” Forced to pick one on IPO day, he takes Anthropic purely on volatility: “more random chaotic turbulent events at OpenAI.”
- Enterprise AI has hit the hangover phase. The sequence: board yells at CEO → token maxing written into performance reviews → “you go and look at the bill and it’s like, oh my god, we are spending so much. I have no idea what the ROI is.” A CIO he knows found hundreds of thousands of dollars per month going to employees asking Opus 4.8 “how’s it going” and what the weather is. Uber’s $1,500-per-individual cap “has happened privately” at dozens of Factory customers, and he expects a short-term contraction in frontier-model usage — which he calls healthy.
- 80–90% of tasks now running on frontier models could run on open models; typically “the planning” needs frontier models. Harry’s pushback: isn’t that “the biggest bear case ever” against Claude Code and Codex? Matan’s answer: the surviving 10–20% are “decision-making tokens” — like leadership hours, few but the most valuable — and per-token frontier spend is rising even as share shrinks. Also: “it’s pretty embarrassing that we don’t have frontier open models in the United States.”
- Token spend is likely comparable to salary. Against Harry’s benchmark (likely Marc Benioff’s $300M Anthropic spend = 3.8% of dev salaries), Matan says the median within three years is “order of magnitude… comparable to salary” — but dispersion runs from 0% to “tens of thousands of percent” per individual, so a standard per-engineer budget is “painting with way too wide a brush.”
- Value accrual is a time-dependent phenomenon — models, apps, and infra are all “trying to commoditize the one that’s not them,” so he strongly disagrees that the next 12 months belong to AI infrastructure. Kirkland’s $500M internal AI build is the counter-lesson: “building AI technology is not a core competency of that firm,” and it’s “actually good for Harvey.” The deeper shift: “there is going to be nothing that no one can build” — moats move from capability to resource allocation.
- Labour displacement rhetoric is fundraising strategy, not forecast. Dario’s “we’re going to take your jobs” line “really upsets me… it’s for selfish reasons” — when raising hundreds of billions, “the best way to convince people is to say all of capitalism is gone” — and it will flip at IPO when those same humans are the buyers. Harry’s addendum: likely Zuck and Demis, who never needed the money, never said it. Short-term displacement worries him; long-term no. Infra bubble? “Long-term absolutely not. Like not even close.”
- Talent and go-to-market are being repriced together: the Olympiad-funnel résumé is now “kind of anti-signal,” the “age of the polymath is back,” and treating sales/marketing as dirty work is a time bomb — AI companies coasting on the gold rush are “astronauts in space where there’s no gravity. Your muscles will atrophy. Gravity will come back.”
Deep dive
1. No single model-war winner — at least four labs comparable
- His most important reversal of the past 12 months: “There was a brief period of time where I thought it might be just one or two companies that run away with being the frontier. What seems pretty clear to me is it’s probably going to be at least four that are probably approximately as good — and that is the win for humanity.” He flags it as a hot take precisely because “right now people are a little bit enamored with maybe one or two.”
- The bear case against Factory is the same trade inverted: “if one model provider gets significantly better than all of the others,” enterprises just go all-in with that lab — “but then that’s a monopoly for the entire economy to be worried about.” His base case is models stay roughly comparable, trading the lead weekly: “one is a little bit better at review, one is better at Python… it all kind of fluctuates every week.”
- Model releases will stop being events: “eventually they’re just going to not announce it and it’s just, here’s our model that’s continuously getting better.” Enterprise engineers “can’t keep up with every single model that comes out, nor should they” — which is, in his telling, the entire case for the application layer arbitraging the cost–quality–speed trade-off per task.
- Quickfire, one IPO-day ticket: Anthropic over OpenAI — the businesses are “approximately equivalent,” so the tiebreak is volatility: “past is an indicator of the future and there’s just been more random chaotic turbulent events at OpenAI.”
2. Everyone is trying to commoditize the one that’s not them
- Against Brendan’s call, likely from Mercor, that the next 12 months are the most value-accruing for AI infrastructure: “I’d pretty strongly disagree.” His frame is the Microsoft org-chart meme with guns pointed at each other — models, applications, and infra each insisting the others are irrelevant. “The reality is value accrual is a time-dependent phenomenon… it’s not like there is one person whose steady state gets all of the value.” He notes everyone talks their book — companies with proprietary data need models to win; Factory needs parity.
- The consumer-optimal structure is models separate from applications, because incentives are misaligned otherwise: “If I’m a model provider giving you a coding tool, I want you to use as many tokens as possible because I’m an API business… I don’t have a huge incentive to be more token efficient.” An independent application layer means “you better damn well be the best or the cheapest or the fastest or else you’ll never get tokens through to you.”
- This may avoid cloud-style lock-in, because of cloud’s scars: sign the three-year discount, “then they would jack up the prices and once you’re standardized on one, it’s going to take you 2 years to switch… you’re stuck with us.” “Every CIO I speak to is really keenly aware of: we cannot throw our lot in with just one model provider.” The same scar tissue produced his most surprising customer: EY, one of Factory’s largest — “they saw scars of being late” to cloud and are now “more agent native than some startups, which is wild.”
- On Lovable/Replit after OpenAI’s competing launch: genuinely unsure — the defensible niche is selling to sales, marketing, and support teams, but chasing non-technical people writing production code would be “ill-advised,” since anything touching databases and access controls “is going to be run by the engineers.”
3. Token maxing is over — the hangover and ROI reckoning are here
- Three phases of enterprise adoption: phase one, “board yells at CEO — hey, what’s your AI strategy?”; phase two, token maxing — adoption written into performance reviews, “the debauchery, the long night, taking shots, using all the AI”; phase three, “the hangover where you go and look at the bill… I have no idea what the ROI is. Is this helping our business?” That’s where most enterprises sit right now.
- His true story from a CIO: “hundreds of thousands of dollars per month” was going to people asking Opus 4.8 “hey, how’s it going? What are my macros from the food I ate today?” — “we don’t need the frontier of human intelligence to be doing this stuff for us. Let alone it’s not even work-related.” Hence routing.
- Uber’s public $1,500-per-individual budget “has happened privately with a lot of customers of ours”: usage exploded before anyone decided which parts of the codebase deserved the tokens. Factory now proactively sets user limits. The consequence: “we might see a short-term contraction of usage of the very frontier models — but I think it’s healthy,” versus staying blind and correcting suddenly.
- Harry’s benchmark question — likely Marc Benioff’s $300M Anthropic spend is 3.8% of dev salaries; what’s that number in three years? Matan: the median is probably “order of magnitude… comparable to salary,” but ranges from 0% for some individuals to “tens of thousands of percent” for those “delegating to dozens of droids in parallel.” An org-wide standard ratio is “painting with way too wide a brush.”
4. Open models can take 80–90% of tasks — frontier keeps the decisions
- His number: “probably 80 to 90%” of tasks currently on frontier models could run on open models — “typically the planning that really needs the frontier models.” Open source is “a really important counterbalance” keeping OpenAI, Anthropic, Google, and Microsoft “under pressure to give the best models as cheap and quick as they can.”
- Harry’s pushback — worth keeping: “Does that not just present the biggest bear case ever against Codex or Claude Code?” Matan’s rebuttal: the 10–20% that stays frontier are “decision-making tokens” — same as human orgs, where “most human hours are not spent on making the decisions,” yet the people making the irreversible strategy calls “are also typically paid a lot.” Frontier spend per token is rising (“ultra high reasoning”) even as its share shrinks.
- There’s an ego trap in the migration: “oh no, the work that I’m doing, only a frontier model could handle” — he admits falling for it himself when switching: “it’s like, no, it probably can.”
- On US startups running on Chinese open models: fine. The sleeper-agent fear — a trigger word that flips the model adversarial — fails a game-theory test: you’d plant it “as late as possible, because if you do that in an early model and someone discovers it, they’re literally never going to use your models ever again.” Still: “I’m quite patriotic. I think it’s pretty embarrassing that we don’t have frontier open models in the United States.”
5. Core competency is the new moat when anyone can build anything
- On whether AI lifts GDP above its 200-year 2% average: “yes, absolutely” — but slowly, because companies are organized around solving problems and “it takes time for the resource allocation to adjust.” Every business now chooses: solve more problems with the same people, or the same problems with fewer.
- The next 24 months’ C-suite question is allocating “dollars, tokens, people” against core competency. Kirkland committing $500M over five years to build in-house AI is his cautionary case: “building AI technology is not a core competency of that firm… I actually think this is good for Harvey — there’s nothing like trying to do something yourself to make you realize, oh shit, this is actually really difficult.”
- The deeper shift: “the world going forward, there is going to be nothing that no one can build. Every single piece of software anyone will in theory be able to build.” His analogy: he knows how to fetch lunch for the team, but “at Factory, our core competency is not that the CEO goes and gets lunch for everyone” — capability stops being the moat; allocation discipline is.
- Org bloat, he argues, came from intermediate metrics: “We shipped four features. What a great quarter. That doesn’t necessarily matter for the business at all.” Same lens on grind-culture theater: “imagine trying to measure who won a basketball game by who sweat the most… look at the scoreboard.”
6. The loadbearing polymath replaces the credentialed engineer
- On likely Andrej Karpathy’s 100x-engineer bifurcation: directionally yes, but 10x of what? “Now I can write a billion lines of code with these tools. It might be shit lines of code though.” His measure is “loadbearing individuals” — remove them and things fall — and AI hands them more leverage while making those who can’t use leverage “that much less valuable on a comparative basis.”
- Harry’s defense of VC credentialism — amid “intense uncertainty around what Anthropic and OpenAI will do and who they will kill,” validators like Olympiad medals are a crutch — is accepted as a crutch. But Matan inverts it for hiring: the high-school Olympiad funnel is “kind of anti-signal, because you’re not owning your fate or choosing your agency,” while the kid from “the middle of nowhere” who sought out competitions alone still carries signal. What matters: “What have you built? How have you taken ownership end to end?”
- The coming role is an ex-engineer GM owning a business outcome — marketing copy, product metrics, sales enablement. And “the age of the polymath is back”: Da Vinci and Euler could reach multiple frontiers because fields were shallow; string theory required “50 years catching up on the literature”; AI shortens that ramp, so he hires people who can hold uncertainty and push several frontiers simultaneously.
- The bottleneck AI won’t solve: “the human side of it… behavior change.” A funny asymmetry — 30-year engineers resist the tools but “know how to delegate,” while early-career engineers adopt eagerly but “don’t know how to manage people.”
7. Sales and marketing are first class — gravity is coming back
- His most controversial opinion in the space rejects the “Silicon Valley fallacy” that research is the pinnacle and sales “dirty stuff”: “The product at Factory is the entire journey from the very first time they hear our name till their 10th renewal after a decade.” Engineers and salespeople sit intermixed: “when salespeople close a deal, engineers say we closed a deal.”
- The warning to gold-rush competitors: “they’re astronauts in space where there’s no gravity. Your muscles will atrophy. Gravity will come back… and you won’t be able to compete.” His gauntlet: “Name a legendary company that has a shit sales or marketing team. You can’t.” Harry’s flip, conceded: “I can name companies that have shit products but great sales and marketing teams.”
- Team-building over grind theater: no mandated hours, no beds in the office — “if you can get your job done on 2 hours of sleep, you’re not doing very high leverage work” — but he bought all 30 employees $3,000 Eight Sleeps during a two-week surge: “it’s like Seal Team Six, the NBA All-Stars — it is worth every dollar to make them more productive.” His answer to Stebbings’ limoncello hedonism: athletes have in-season and off-season — locked in, then (Harry’s line) “out of season, you Charlie Sheen.”
- His enterprise-sales education, from someone whose first job ever is this one: face-to-face matters enormously, and “you should never try to sell something” — understand problems and check fit. On FTEs in deals: they should only accelerate consumption (“scale to a million dollars in three months instead of six”) — “if we need FTEs to make the product work, we have a shit product. I’m not Accenture here.”
8. Engineers stop writing code and start building the factory
- Rollout phase one was “look how much code we can generate”; phase two, “some poor staff engineer who has to review hundreds of these slop PRs.” The fix is agent-readiness — up-to-date docs, remote machines where agents run their own code, CI/CD, linters, pre-commit hooks — whose ROI used to scale 1:1 with engineer count and is now “10x or 100x depending on how many agents you’re using.”
- The company name is the thesis: organizations will have “engineers that build the factories that build their software” — his image is Tesla’s assembly line, few humans in it, humans designing all of it to maximize throughput and stop agents “technically passing tests but dramatically increasing debt.”
- What we’ll laugh at in five years: highly paid engineers hand-writing release notes and documentation. Stripe’s famous docs edge gets equalized — “yes” that reduces its impact, but “it’s a better world where everyone has documentation as good as Stripe’s.”
- The lag to watch: code generation is “growing exponentially. The security efforts aren’t growing in kind… there are probably going to be in the next couple years some pretty big incidents” — likely already happening unadmitted — “and we haven’t even seen the most adversarial behavior yet.”
9. Labour displacement fear is fundraising rhetoric — and Dario did damage
- On displacement: “short-term, yes” it worries him — the layoffs are a shock hitting tens of thousands. “Long-term, no”: “there is a huge number of problems in the world… and very few of those problems that can be solved with software are we currently solving with software.” Flooding the market with engineers redistributes talent to health, pharma, climate — though “the economy has to match and properly incentivize them,” which takes time.
- His sharpest words go to the pause-AI camp, via dementia: “that is something that can be solved with better AI and better software… by saying you want to slow down AI, you’re saying to people with loved ones who have dementia: no, no, sorry, you’ve got to maintain that relationship a little bit longer. We’re scared.” His verdict: “pretty harmful and pretty selfish.”
- Asked directly whether Dario did the ecosystem a disservice with “we’re going to take your jobs”: “Yes… not only disingenuine and wrong, but it’s really hurt the psychology of a lot of people — and it’s for selfish reasons. If you’re trying to raise hundreds of billions, the best way is to say all of capitalism is gone, the only company left will be me. Then when it comes to IPO… suddenly it’s whoa, whoa, whoa — humans are pretty important.” Harry’s addendum: the ones who never said it — likely Zuck and Demis — never needed the money.
- On government steering of AI talent: reluctant beyond defense and safety — “you need to have a very good case” — and even climate may argue for faster AI despite near-term emissions. AI infrastructure bubble? “Maybe some short-term blips, but long-term absolutely not. Like not even close.”
10. From string theory to a $1M Sequoia check at 20% post
- The origin: a geometry teacher told him to retake the class, so his first-ever Amazon order was a stack of math textbooks he worked through in one summer, placing out of everything. His dad said the hardest math was string theory — cue 12 years of tunnel vision: Princeton (first undergrad to work with and write a paper with likely Juan Maldacena), then a Berkeley PhD where “everything came crashing” once he realized he’d been doing it “because it’s hard and because someone said I couldn’t.”
- A program-synthesis seminar “completely nerd sniped me”: “code with the explicit purpose of creating itself” — physicists never care about n=3, only the n-dimensional solution. He read Zero to One, then heard an ex-physicist Sequoia partner (likely Shaun Maguire) he’d once cited, on a podcast. A cold email became a three-hour Sand Hill walk; he withheld the pitch — “didn’t want to dirty it” with a transaction.
- The next day he met co-founder Eno at a hackathon (“intellectual love at first sight”), they rebuilt the demo in 72 hours, and the partner said “drop out of your PhD and send me a screenshot” — brutal for parents who’d emigrated from the Soviet Union with nothing. Partnership pitch in April 2023, pre-Copilot adoption, pitching fully autonomous agents: $1M at 20% post. Harry’s math: that stake is roughly a $300M position at the last round, pre-dilution.
- Why he never shopped it: “no one else would have believed in me… trust and loyalty and belief matter so much more than the price tag.” Take a discount for Sequoia? “Generally yes” — what counts is “deep conviction when it’s not obvious,” not hot-round flattery, which he learned firsthand from an old-guard investor: “I’m the fucking man… and 30 minutes after it wore off — oh my god, he got me.” On Ivanka Trump (on the cap table via his hire Francesca and the Affinity connection): genuinely valuable — “there is dirty work, investor help that she does that some investors who are more known as investors do not do.”
Verification Notes
- The raw captions contain both “Five post” and later “20% post”; the digest retains the later figure.