The Supply and Demand of AI Tokens | Dylan Patel Interview
The Supply and Demand of AI Tokens | Dylan Patel Interview
Summary
- SemiAnalysis is living the demand explosion it covers: AI spend was “tens of thousands” last year, rising to a $7M annual run rate on Claude Code (it was $5M just a week before taping) against ~$25M of salary — north of 25% of payroll and, on trajectory, “more than 100% by the end of the year.” Dylan’s framing is existential, not optional: “if I don’t adopt AI, someone else will and they will beat me.”
- The Anthropic margin math is the episode’s most tradeable number: revenue went from $9B to “$35, 40 billion” ARR while compute didn’t grow proportionally, implying a gross margin floor of 72% versus “30-something percent” in leaked funding docs at the start of the year. Demand is so far above supply that Anthropic could “double their pricing on Opus and I would continue to pay… I bet that wouldn’t solve their humongous capacity problem.”
- Claude Mythos is “potentially the biggest step up in model capabilities in like 2 years” — Anthropic went from an L4-equivalent software engineer (4.6 Opus) to an L6-equivalent (Mythos, internally available in February) in two months, and is withholding it except for selective cybersecurity use at “five or 10x the token cost”; top banks have access, while Anthropic ships Opus 4.7 “preferentially made worse at Cyber” per its own model card. Narrowing model deployment, not broad access, is the new regime.
- The structural inversion: “ideas are cheap and plentiful, but execution is very easy” — which is why lab release cadence compressed from 6 months to 2, and why pure diffusion of a 4.6-Opus-tier model takes economy-wide spend from $40B to $100B by year-end on a linear extrapolation alone. Mythos if the world had enough compute “would be $500 billion of revenue or something crazy” — so even the tier-two lab sells out of tokens and the tier-three lab is probably close to sold out, and lab margins expand until the hardware and infrastructure supply chains jack their own.
- The distributional call, delivered unhedged: “If you don’t use more tokens, you’ll never escape the permanent underclass” — you must use tokens, generate value from them, and capture that value. He raises concentration risk: he sketches Ken Griffin buying “the first $10 billion of tokens each year” on every new model release and crushing the market.
- Supply is sold out everywhere: DRAM “will double or triple from here” with true incremental capacity not arriving until ‘28; GPU useful life is stretching to “maybe even 7 or 8 years” as Hopper and A100 clusters re-sign at higher prices; TSMC may “sincerely” spend $100B on capex in 2028; CPUs are “completely sold out” on RL environments and deployed AI code; next-gen AI racks carry 120 FPGAs each. Wafer-fab equipment (Lam, Applied, ASML, MKSI) is “still very underappreciated.”
- His three-month prediction is categorical: “large-scale protests” against Anthropic and OpenAI — AI polls “less popular than ice, less popular than politicians,” and Sam Altman had “a Molotov cocktail thrown at his house twice in like 2 weeks” to cheering comment sections. His fix: Sam and Dario “have to stop getting on interviews. They’re so uncharismatic.” Meanwhile robotics few-shot breakthroughs in 6-18 months could extend token demand.
Deep dive
1. Claude Code is 25% of SemiAnalysis payroll — and tracking past 100%
- Last year “tens of thousands of dollars” felt like heavy AI usage. Then Opus landed in late December, and president Doug O’Laughlin was “very much… leading the charge” among nontechnical people using AI for coding; he “pulled the whole firm slowly over time,” and spend inflected into an enterprise Anthropic contract and a $7M annual run rate on Claude Code ($5M when Patrick last saw him, a week earlier) against ~$25M of salary. “If this trajectory continues, we’ll spend more than 100% by the end of the year — which is a bit terrifying.”
- Dylan dodges the people-versus-AI choice only because “our company’s growing so fast.” Others won’t: “if this person can do the work of five to 10 to 15 people using Claude Code, then all of a sudden I should probably cut people.”
- The specimen example: SemiAnalysis’s Oregon reverse-engineering lab. One ex-Intel employee spent “a couple thousand dollars of Claude tokens” building a GPU-accelerated app that overlays every material on a chip image — copper, tantalum, germanium, cobalt — enabling fast finite element analysis of the whole stack-up. At Intel, “that was an entire team’s job to build and maintain.”
- Malcolm, formerly at a bank with a 100-200 person economist department, alone built what “would have taken the team of 200 economists a year”: piped FRED and employment data, graded the BLS’s ~2,000 tasks for AI-doability (~3% doable now), and coined “phantom GDP” — output rises but costs fall so far that “GDP theoretically shrinks.” “He’s just completely cracked out on Claude.”
2. Adopt AI or be commoditized — the energy-grid proof
- The business logic: SemiAnalysis sells information, and “I don’t see why this wouldn’t be completely commoditized on a pretty rapid basis if I’m not constantly improving” — his 2023 flagship dataset “is basically what everyone else is doing now.” Hence the existential version: “if I don’t adopt AI, someone else will and they will beat me.”
- A year of multiple energy analysts failed to crack the ~$900M energy data-services market. Then “Claude Code psychosis” hit Jeremy, who leads data-center energy: spending ~$6,000 a day, he scraped every US power plant and every transmission line above a certain voltage into a dashboard of micro-regional power deficits and surpluses, building it in three weeks. Energy traders’ verdict: “better than XYZ company” — a firm with 100 people and a decade of work. “Who’s going to come commoditize me if I don’t move faster?”
- Patrick’s pushback — won’t capital-rich funds just build this in-house? Dylan’s answer: information always sells below the value it creates (“if I sell you information for a dollar, you’re only buying it because it lets you make more than $1”), and even the Jane Streets and Citadels keep buying because buying and building on top is cheaper than building. “Some may try.”
3. Anthropic’s margins: from a leaked ~30% to a 72% floor
- The macro math: Anthropic went from $9B revenue to “$35, 40 billion now… probably 40, 45 by the time this airs.” Compute hasn’t grown to the same degree, and research compute clearly wasn’t cut (they shipped Mythos and Opus 4.7) — so even assuming all incremental compute went to inference, gross margins floor at 72%, versus “30-something percent” in funding-round docs leaked at the start of the year.
- How margins expand like that: demand so high Anthropic can cut usage and rate limits. “What really matters is having an Anthropic rep and an enterprise contract” — and his advice verbatim: “Whoever you are, if you have enough capital, you should get a freaking enterprise Anthropic subscription where you pay per token.”
- His pricing stress test: “they could double their pricing on Opus and I would continue to pay and I bet most users would continue to pay. I bet that wouldn’t solve their humongous capacity problem.”
4. Mythos: L4 to L6 in two months — and they’re sitting on it
- His funniest recent memory: he and his buddy Leopold “on our knees in front of an Anthropic co-founder begging him for access to Mythos” while the co-founder pretended it didn’t exist. On the benchmarks, Mythos is “potentially the biggest step up in model capabilities in like 2 years.”
- Anthropic’s stated goal was an L4 software engineer by end of 2025, achieved with 4.6 Opus. Mythos benchmarks “like an L6 engineer” and was internally available in February — L4 to L6 in two months. “What’s next?”
- The release pattern is the story: Mythos had a selective cybersecurity release at “five or 10x the token cost”; top banks have it for cybersecurity, while the public got Opus 4.7, “a shitty worse version,” whose model card admits “we actually preferentially made it worse at Cyber.”
- On unit economics, Mythos costs more per token but “spends a lot less tokens to do the thing,” making it “actually cheaper in most tasks than 4.6 Opus.” It’s also a materially larger model — “proof that the scaling laws still work” — with Anthropic going for “a huge jump” where OpenAI scales “in small steps.”
5. Ideas are cheap, execution is easy — the economy reorders
- The episode’s spine, in his words: “What used to matter was execution was very, very difficult and ideas were cheap. Now, ideas are cheap and plentiful, but execution is very easy. So really only the good ideas can justify the spend on super cheap implementation.” Applied to the labs themselves, that’s why “release cadence has shrunk down to 2 months from where it was 6 months before.”
- He says uncertainty is there, but fears “how does society reform itself” when implementation ability stops mattering — what counts becomes choosing the right idea, selling it, and garnering capital toward it.
- Falling cost curves aren’t the demand driver: GPT-4-class capability fell to 1/600th the cost (DeepSeek) and further, but “no one gives a crap about GPT-4 class models” — the frontier creates the economic value. His own spend at today’s quality would be ~$70k in a year, “irrelevant, because I’m going to be using a way, way better model.” Patrick’s confession proves the point: rate-limited mid-flight the day 4.7 launched, “I couldn’t think about using 4.6 anymore.”
6. Even the tier-two lab sells out — and robotics is the second curve
- To “Anthropic’s just won, OpenAI’s cooked,” Dylan says no. Anthropic is compute-bound — “Dario used to gloat about how OpenAI was being too aggressive on compute, and now Anthropic is like, I wish we had a lot more compute” — while OpenAI raises money against “irresponsible levels of compute” from Oracle, CoreWeave, SoftBank, Microsoft, even Trainium from Amazon, ahead of its alleged [spud?] release (per The Information).
- The diffusion math: pure adoption of a 4.6-Opus-tier model takes economy-wide spend from $40B to “$100 billion by the end of the year” — “a linear extrapolation, not an exponential.” Mythos with enough compute “would be $500 billion of revenue or something crazy.” So “even the tier-two lab is going to be sold out of tokens… probably the tier-three lab will also be close to sold out,” and lab margins keep expanding “until people in the hardware supply chain are like, wait, why don’t I just jack up my margins?”
- Robots consume roughly zero tokens today, but he calls the software-only singularity “just a blip.” Current VLAs (vision language action models) are data-inefficient and “probably not going to be the thing that ultimately scales,” but with implementation cheap he expects few-shot robot learning breakthroughs in 6 to 18 months — pre-trained robot models you show a few examples, then niche downloadable skills (“robots just for cleaning chalkboards”). “I don’t think token demand slows down, personally.”
7. Use, generate, capture — or join the “permanent underclass”
- The line of the episode, delivered without hedge: “If you don’t use more tokens, you’ll never escape the permanent underclass.” He splits it into three distinct problems — using more tokens, generating value from them, capturing that value. The “boring lazy way” is working one hour instead of eight; “the cool way is I’ll still work eight hours a day and do 8x the work and maybe make 5x the money.”
- Broad deployment is narrowing: labs fear distillation and real-world impact, so “models will have less broad and less broad deployment” — he trolls Anthropic by calling the project “earwig,” though he says that isn’t its name. His hypothetical: Ken Griffin signs for “the first $10 billion of tokens each year” on every model release — “now he’s going to crush everyone in the market.”
- Nobody knows the capability frontier — “Anthropic doesn’t know what these models can do. No one knows” — so leverage falls to end users, which is “tremendously productive and uplifting for humanity, but then what happens to the concentration of resources?” Maybe a year or two from now, “the business is actually just arbitraging tokens.”
8. Supply: anything “that has a pulse” is sold out
- The GPU-depreciation bears are wrong: three-to-four-year-old Hopper clusters are re-signing for three or four more years, A100s for a couple more — useful life is “maybe even 7 or 8 years,” not under five, with prices rising on renewal, so cluster gross margins beat the assumed 35%.
- The sharpest call is memory: capacity grows only 20-30% a year and true incremental supply “doesn’t come till ‘28, late ‘27 at best” — so “DRAM will double or triple from here” still, because capacity must be stolen via demand destruction: “we’re not rationing stuff here.” People who say the memory story is overplayed “don’t get it.”
- TSMC says this year’s CapEx is $56B; Dylan says, “We’ve had $57.4 billion since January,” and “sincerely, they may spend $100 billion on CapEx in 2028 — people just can’t fathom that.” Yet TSMC takes single-digit price increases “because they’re good people, it seems like,” while wafer-fab equipment — Lam Research, Applied Materials, ASML, MKSI — stays “very underappreciated” as the capex tail “gets whipped harder and harder.” ASML is sold out and “needs Carl Zeiss to expand faster”; down to copper foil and glass fibers for PCBs, sold-out inputs attract prepayments that lift return on invested capital even where gross margin doesn’t.
- The overlooked compute layer: reinforcement-learning environments and all the deployed AI-generated code running on what is likely a Vercel instance or an AWS instance sit on CPUs, which are “completely sold out” — and a next-generation AI rack carries 120 FPGAs.
9. The number he can’t compute — and the protests he predicts
- What he wishes he knew: tokenomics on the usage side. Infrastructure costs and lab margins he models well, but adoption keeps breaking his models — “In January, we had crazy estimates for February. Anthropic smashed them,” then again in March. “Who is using all these tokens? What are they building with them?”
- The deeper measurement problem loops back to phantom GDP: token value diffuses as better decisions across the economy and “it’s not really something you can capture in any GDP statistic.” “It’s clearly by every subjective metric amazing. But where is the phantom GDP?”
- His three-month prediction, stated flatly: “Large-scale protests” against Anthropic and OpenAI. “AI is less popular than ice, less popular than politicians — confused how Pew surveyed this.” Sam Altman “had a Molotov cocktail thrown at his house twice in like 2 weeks” and news-article comments cheered. “This is just the beginning.”
- The counterweight, undiplomatic: “Sam Altman and Dario have to stop getting on interviews. They’re so uncharismatic… Sam being on Tucker Carlson probably made all Republicans hate OpenAI.” Show present-day uplift, stop preaching future world-changing capability to people who see the labs as “a sneaky cabal of 5,000 people” — “a huge reorg and rebranding needs to be done.”