Pioneers Insight Method Research Author
Ep. 013 - AWS Margins Jump 10% While Azure and GCP Flatline
Back to Episodes

Ep. 013 - AWS Margins Jump 10% While Azure and GCP Flatline

Summary

  • AWS’s improving margins versus Azure and GCP come from selling Claude through Bedrock as higher-margin tokens, not simply renting accelerators. AWS is adding more than a gigawatt of capacity per quarter while margins rise; Azure and GCP remain far more exposed to lower-margin infrastructure-as-a-service. Joey’s framing: token sales retain more upside than five-year take-or-pay contracts.

  • The capacity ramp normally crushes near-term cloud margins before clusters reach “stabilized” utilization. CoreWeave-style providers pay depreciation, leases, and labor while complex systems such as GB200 wait months for activation and produce no revenue. AWS’s ability to absorb the same costs while expanding margins is therefore “a pretty good sign” for its eventual return on capital.

  • Claude’s API-heavy growth is giving both Anthropic and AWS unusually strong operating leverage. Joey cites Anthropic at $47 billion of ARR, with probably $10 billion of net-new ARR per month across March, April, and May and roughly 80% of the increase coming from APIs. Amazon was “in the right place at the right time”: Claude represents an estimated 80%-93% of Bedrock usage.

  • Anthropic’s $65 billion Series H at a $965 billion post-money valuation looks less extreme against its growth and profitability. Joey compares the roughly 20x ARR multiple with the 80x levels reached by software names in 2021 and says Anthropic is profitable excluding stock-based compensation. The caveat is significant operating deleverage if enterprises curb coding-token consumption, but “there’s no train that’s slowing right now.”

  • The SpaceX/xAI compute deal produced the episode’s sharpest disagreement over AI demand. Jeremie sees a former compute buyer becoming a supplier and possibly “giving up on the frontier race”; Jordan sees overwhelming Anthropic demand, valuable GPU-recall optionality, and a rational way for xAI to earn revenue until its own research and distribution can use the capacity.

  • Whether AI becomes winner-take-all depends on whether spending concentrates in open-ended tasks where “good enough” never arrives. Crystal argues a third- or fifth-ranked model could still replace substantial labor; Jeremie counters that legal work, science, healthcare, and analysis reward continually buying more intelligence. Jordan’s formulation is even broader: “coding is not coding, it’s computer use.”

  • The durable winners may be hyperscalers combining frontier-model access, enterprise distribution, and custom silicon. Bedrock could become the majority of AWS’s AI business by year-end, while Azure and GCP remain 80%-90% infrastructure-as-a-service in the panel’s model. Trainium and TPUs gain another advantage when token buyers never need to know which accelerator served them: “Winners win, losers lose.”

Deep dive

1. Bedrock turns cloud capacity into a higher-margin product

  • Joey divides hyperscaler AI into three businesses: software such as GitHub Copilot, infrastructure-as-a-service that rents accelerators, and “token as a service,” where customers buy model access through existing cloud agreements. The last retains enterprise security, availability zones, and consolidated billing while letting the cloud capture more economics than bare chip rental.

  • Crystal argues that GPU-as-a-service lowered the old cloud moat: major users increasingly want “just the metal,” configured their way, rather than the managed platform that once made AWS, Azure, and GCP exceptionally defensible. Jordan says there are now more than 200 neo-clouds, illustrating the lower barriers to entry.

  • Bedrock restores a differentiated economic profile: Claude usage routes predominantly through AWS, and token sales carry better margins than renting infrastructure under a fixed contract. That mix explains why AWS operating margins are improving while Microsoft’s decline and Google’s remain comparatively flat.

2. AWS is outrunning the capacity-ramp margin trap

  • Crystal’s CoreWeave example separates steady-state economics from the ramp: a fully functional, stabilized cluster might produce roughly 25% operating margins and 30%-40% gross margins under a five-year take-or-pay contract. The operative word is “stabilized”—contracted revenue and the stable cost profile apply only after the cluster is functioning.

  • Before activation, the provider has already built the data center and begun paying depreciation, leases, labor, and other expenses. Equipment installation takes months, and complex systems such as GB200 have lengthened that interval, leaving assets generating costs but no revenue.

  • AWS faces the same physical constraints while bringing on more than a gigawatt per quarter, yet its margins are rising. Crystal’s inference is explicitly conditional: if margins expand during unprecedented delivery, the eventual stabilized margin and return on capital for token-as-a-service could be “extremely rich.”

3. Claude’s API mix makes Anthropic unusually measurable

  • Joey estimates Claude accounts for 80%-93% of Bedrock, while Microsoft skews toward OpenAI and Google toward Gemini. He cites Anthropic at $47 billion of ARR and probably $10 billion of net-new ARR per month across March-May, with about 80% of that incremental ARR coming from APIs.

  • Crystal finds Anthropic easier to forecast than OpenAI because API workflows expose token consumption and pricing has stayed relatively stable. OpenAI’s Q1 business was roughly 60% consumer subscriptions, where users can cancel or switch; Anthropic’s API consumption is more observable.

  • On Opus 4.8, Crystal preserves the caveat: regular pricing is unchanged, fast-mode pricing differs, and Anthropic’s chart “supposedly” shows substantially fewer hallucinations than 4.7. She had not tested it, but hoped it would stop “pulling numbers out of thin air” in their analyses and said it was close to Mido’s preview.

  • The $65 billion Series H values Anthropic at $965 billion post-money, nearly double the roughly $380-$400 billion figure discussed for February. Joey calls the approximately 20x ARR multiple less extreme than 2021 software valuations and says Anthropic is profitable excluding stock-based compensation, though coding-token cuts could cause significant operating deleverage.

4. The SpaceX/xAI deal divides the panel on compute scarcity

  • Jordan says the SpaceX filing revealed “many billions” in spending with SpaceX, with a provision allowing xAI to reclaim the GPUs. For Anthropic, the capacity could relax rate limits if compute constraints were limiting growth; for xAI, it converts capacity into a revenue-producing asset without permanently surrendering it.

  • Jeremie reads the deal as bearish for aggregate compute demand: “one player that was supposed to be a source of demand becomes a source of supply,” increasing GPU-as-a-service competition while removing an offtaker. More starkly, selling capacity to Anthropic suggests xAI may be abandoning the frontier despite evidence of extraordinary returns to training.

  • Jordan’s rebuttal is that the deal exists because Anthropic’s demand is “so overwhelming” that it must buy capacity from competitors. The recall clause also matters: xAI can monetize compute now, do research on the side, and redirect the fleet after a breakthrough or once distribution through X, Starlink, or Tesla can support a larger run.

  • The disagreement becomes capital-structure specific. Meta can fund long-duration research from a huge cash-generating business; xAI relies on finite venture capital. Joey’s “earnings before training interest and taxes” framework treats inference profit as operating cash and training as investment, raising the question of whether each lab’s model spending earns an adequate return.

5. Frontier economics may leave little value below the leaders

  • Crystal asks whether ranking even matters: a third-, fourth-, or fifth-best model might still be good enough to replace many jobs. Jordan defines the race differently—“coding is not coding, it’s computer use”—and ranks Anthropic first, OpenAI second, and Cursor third because Composer combines distribution with a usable model.

  • Jeremie rejects “good enough” for the largest pools of spending. Translation may eventually become finite, but legal work, science, healthcare, and analysis are open-ended: users can spend more intelligence to gather evidence, explore alternatives, and beat competitors. His conclusion is categorical: “If you’re three or four or five, you’re not gonna get any dollars.”

  • His Mythos example distinguishes token price from completed-task cost: he calls it “two-thirds cheaper than Opus” despite token pricing being “six times, I think, five times more expensive.” If a model is 10x smarter and needs 10x fewer tokens, it might still be the cheaper tool—concentrating demand at the frontier.

  • Compute alone does not settle the race. Jordan argues that labs need both compute and talent; Jeremie agrees, and both see xAI’s loss of talent alongside its compute as evidence that a comeback will be difficult. Jordan remains unsure which combination of revenue, distribution, and “hero runs” ultimately wins.

6. Distribution and silicon keep the market concentrated

  • Jordan challenges the winner-take-all thesis with Cursor at roughly $2 billion of ARR and Fireworks claiming $315 million. Jeremie concedes encouraging open-source adoption but calls those figures immaterial beside an AI market already above $100 billion: “What is $300 million between friends?”

  • Joey expects Bedrock could become the majority of AWS’s AI business by year-end, even though AI remains a smaller share of AWS than of Azure or GCP. Azure and GCP are still estimated at 80%-90% infrastructure-as-a-service, but Azure could add Claude quickly because hyperscaler customer relationships make token distribution relatively easy.

  • Neo clouds face a tougher loop: successful token-as-a-service requires frontier-lab partnerships, speculative capacity, and enough capital to deploy GPUs without a five-year offtake. CoreWeave, Nebius, and AIREN lack some combination of those advantages, leaving inference endpoints overwhelmingly concentrated among the top three hyperscalers.

  • Custom silicon widens the margin gap. Trainium at AWS and TPUs at GCP provide vertical integration that NVIDIA-heavy Azure lacks; usability matters less when a gigawatt serves only a few models. As Jordan puts it, a Claude Code user cannot tell whether a token came from a GPU, TPU, or Trainium—and does not need to.