Ep. 020 - Anthropic vs OpenAI Usage, Margins, Meta Compute, Future of MSL (Tokenomics)
Summary
Enterprise token austerity is aimed at the wrong workloads: coding consumes the budget, while premium-model emails are effectively free. Crystal says most employees never approach their caps, so pooled company budgets make more sense than per-person ceilings. Max calls engineering-only access a “caste system,” while Joey says the top 1% of companies already spend roughly $100,000 per employee annually on AI and continue when the ROI survives scrutiny.
The $200 coding subscriptions are heavily subsidized and may be loss-making at power-user utilization. The team estimates Codex can provide about $12,000 of API-equivalent usage and Anthropic about $8,000; break-even utilization falls to 10% for Claude Max 20x and 5.7% for OpenAI Pro 20x. At likely utilization levels, Max says these plans “might just be negative margin.”
Anthropic’s enterprise/API-heavy mix is already producing operating leverage that OpenAI’s free consumer base cannot match. More than 80% of Anthropic ARR is API-based and historically more than 90% enterprise; Joey says some profitability is on a non-GAAP basis excluding stock compensation, while operating profit was positive in Q2 and could reach $1 billion-plus in Q3. OpenAI’s roughly 950 million weekly users convert only about 6% to paid plans, and servicing the free population lowers blended gross margin by around 20 points.
OpenAI has nevertheless returned to a genuine two-horse race because model quality, not incumbency, remains decisive. After likely flat ARR growth in March and April, 5.5 became the predicted inflection point; some people now judge 5.6 “as good as Opus 4.8” at half the price and likely with a smaller model. Net-new monthly ARR is reportedly comparable with Anthropic, prompting Dylan’s reversal: “I am nothing if not flexible.”
Compute optionality may matter as much as compute ownership for xAI and Meta. Max’s revised view is that surplus capacity can be rented at 3x-4x market rates while teams retain enough to prove they can reach the frontier, protected by 90-day clawback clauses. Meta compute could similarly justify heavier 2027-28 capex and become a profitable fallback if MSL disappoints; by contrast, Max reads Google’s non-recallable, long-term TPU commitments as evidence of weak conviction in building RSI.
Hyperscalers win token-as-a-service distribution through existing enterprise relationships, while independent inference providers can prosper without winning the market. Anthropic’s indirect token volume may have risen from roughly 5% to 20% of its business in six months, and AWS may collect a 20%-30% revenue share largely because customers already buy through Bedrock. Together, Fireworks and Baseten can still grow rapidly because inference could be “the largest market ever,” even if open-source volume grows more slowly than frontier-model volume.
RL environments are becoming both the next capability bottleneck and an unusually lucrative engineering market. Frontier-lab data budgets could exceed $10 billion this year — about 10x last year — and might increase another 10x in 2027; top coding tasks command well over five figures each, while people who are very good at creating RL tasks can earn seven figures or more annually. The episode makes its exponential thesis falsifiable with a $400 billion Anthropic ARR over-under for end-2027: Max and Dylan take the over, Crystal, Joey and Jeremy the under, and the loser must explain publicly why they were wrong.
Deep dive
1. Token caps are targeting the wrong workloads
Crystal’s field read is that social media exaggerates token-budget pain: only a handful of power users at most companies approach the limit. Because usage is highly uneven, one company-wide monthly pool is more rational than identical per-person allowances.
Smaller organizations are already routing around limits: a cheap model processes and condenses material before an expensive Anthropic or OpenAI call. Where internal models do not count toward the budget, employees “really push those to the limit” and reserve metered models for the final work.
Max’s objection is that many policies optimize trivia. Whether a salesperson uses Opus or Sonnet for an email “basically does not matter” from a token perspective; coding is “by far the most token-hungry use case.”
Restricting Claude Code or Codex to engineers creates what Max calls a “caste system,” denying other functions the chance to discover valuable workflows. Joey’s supporting evidence: he estimates that 70% or more of Anthropic’s API-side ARR is in coding, while Meta contributes only about 3%-5% of total Anthropic ARR — usage is heavily driven by power users.
2. Subscription arbitrage exposes the real margin pool
Max’s practical rule is blunt: “Anybody who can use a subscription definitely should,” because it is so heavily subsidized. Some startups simply charge five or however many individual accounts they need to the company card and add another whenever they hit a limit; Max says buying 10 subscriptions per person can be more cost-effective than API credits.
At $200 monthly, the team estimates Codex supplies roughly $12,000 of API-equivalent credits and Anthropic around $8,000. The corresponding break-even utilization is about 20% for the Claude Pro and 5x plans, 10% for Claude Max 20x, 11.4% for OpenAI Plus and Pro 5x, and 5.7% for Pro 20x.
Max believes average 20x-plan utilization exceeds those thresholds, making the products potentially loss-making rather than merely lower-margin. Anthropic API gross margin is estimated at 85%+, and Anthropic enterprise subscriptions include no usage: customers pay the fee, then consume at API prices.
3. Anthropic’s API mix converts model demand into profit
More than 80% of Anthropic ARR is API-based and historically more than 90% comes from enterprise. Joey says some of the company’s profitability is on a non-GAAP basis excluding stock-based compensation, while operating profit was positive in Q2 and could reach $1 billion-plus in Q3.
That profitability partly reflects an inability to hire or reinvest in training as quickly as revenue is growing. Joey’s prescription is not to race toward 30% EBITDA margins, but to spend the gross-profit advantage extending model capability: “continue to invest as much as you can into training.”
OpenAI had the inverse mix in Q1: roughly 60% consumer ARR, around 950 million weekly active users, only 6% paying, and most subscribers on $20 plans. Free users lower blended gross margin by about 20 points; Dylan notes reported daily active-user counts rising from 7 million to 8 million against more than 900 million free weekly users.
The team believes Codex, 5.5 and 5.6 may already have flipped OpenAI to roughly 40% consumer and 60% enterprise in Q2. Its closing wager captures the stakes: Anthropic ARR above or below $400 billion at end-2027, with Max and Dylan over, Crystal, Joey and Jeremy under.
4. Better models have pulled OpenAI back into contention
Crystal says Claude Code initially retained users through habit — “that’s just what I’ve been using” — but ChatGPT’s consumer traction is now trickling into work. Former Claude Code power users who ignored Codex are beginning to experiment with it.
Max believes OpenAI ARR growth was likely flat in March and April, setting off “alarm bells,” before 5.5 delivered the expected inflection at roughly Opus 4.8 quality. Some people now rate 5.6 level with Opus 4.8 despite half the price; the team also believes it is likely much smaller.
Dylan’s change of mind is model-led: he now defaults to OpenAI for his work, but rejects the Codex app’s thread model, weak remote-SSH experience and missing tab workflow. He remains a CLI user because the app appears aimed at global “vibe coders” making more B2B SaaS.
5. Frontier compute is most valuable when it can be recalled
Max dismisses Grok 4.5 and Muse Spark 1.1 as adequate scaling proofs rather than frontier products: “unless you can deliver a true frontier model… who cares.” He plans to shift his token usage to either, though their existence proves a nonzero chance of catching up.
His harsher ranking puts Gemini “clearly in fifth place,” potentially forever unless Gemini 3.5 Pro beats current industry chatter. The distinction is not merely installed compute, but whether management preserves the option to redirect that compute toward frontier training.
Max revised his interpretation of Elon renting nearly all of Colossus 1 to Anthropic. The strategy is to leave the internal team enough capacity to prove scaling works, rent the rest at perhaps 3x-4x market rates, and retain a 90-day clawback if the model reaches the frontier.
Dylan’s pushback — worth keeping — is that a provider may only claw capacity back once before losing customer trust. Max agrees that one recall is ideal, but argues Anthropic might return anyway if Elon later relents and Anthropic remains desperate for capacity: “Why would they say no?”
6. Meta compute is a capex backstop, while distribution powers token resale
Meta can apply the same structure through recallable rentals, token-as-a-service, or capacity assigned to MSL, preserving the option to return everything to MSL. Joey sees this as a backstop permitting more aggressive capex in 2027 and potentially 2028 if MSL is deemed unsuccessful.
Selling one’s own frontier-model tokens remains the best business, but premium compute rental is hardly a consolation prize. The discussed SpaceX-plus-Google transaction priced near 4x market, versus neocloud economics around $12 billion per gigawatt and potentially approaching $50 billion.
Anthropic’s revenue density illustrates why buyers can tolerate those prices: about $16 million per megawatt last year, more than doubled over three quarters, with some expecting another doubling within two or three quarters. “Token throughput’s been pretty impressive,” and new models monetize at higher prices.
Dylan initially read an older chart as evidence that Microsoft Foundry was doing badly; Joey corrected him that it was roughly 90% OpenAI until recently and should be revised upward after recent API gains. Bedrock and Azure remain attractive because Fortune 500 and Global 2000 customers already have large cloud relationships, credits or ELAs, plus security and compliance.
7. RL environments are becoming the next capability scaling constraint
Max calls RL “probably the most important scaling law” for current capability gains. Many people believe sufficient environments could teach models nearly any computer-based human task through repeated attempts; aggregate frontier-lab data budgets may exceed $10 billion this year, about 10x last year, and potentially 10x again in 2027.
This market has almost no demand constraint: labs will not reject excellent data merely because it is expensive. Max attributes Anthropic’s coding lead before 5.6 partly to buying coding environments far more aggressively than rivals, which are now trying to catch up.
Modern environment creation bears little resemblance to bounding boxes or content labels. A strong software task may require a really good human engineer’s full day, integration-test verifiers and a rubric, plus a prompt that is simultaneously unambiguous enough for fair training and natural enough to resemble a real request.
Labs will pay well over five figures for one top coding task, and people who are very good at creating RL tasks can make seven figures or more annually. Max’s own experience is that creating data is much harder now than it was eight months ago: identify a precise failure mode, let the model attempt it ten times, then iterate. “Your first thought is almost certainly too easy — try making it 10 times harder.”