Nebius Co-Founder on AI Infrastructure Bubbles | How Price Elastic is Demand for Compute
Nebius Co-Founder on AI Infrastructure Bubbles | How Price Elastic is Demand for Compute
Summary
- Chernin’s anti-bubble case rests on how early adoption actually is: coding is the only AI use case working at scale, and it started working “maybe a few months ago” — every large company, even technologically advanced ones, is “in the first percent of the volume, the first percent of the use cases.” He concedes bias (“I would probably not be in the business if I wouldn’t believe”), but the enterprise-adoption math alone says “we are only beginning.”
- The single most tradeable anecdote: during the DeepSeek panic, Nebius stock fell 40% in one week — the same week the company had “the best commercial week in the history of the company.” Cheaper intelligence didn’t cut consumption; it made previously uneconomic inference workloads viable and lit up Cursor’s growth. “Every time we got the same unit of intelligence cheaper, we are not reducing the consumption but increasing the consumption.”
- Demand is real and unfilled at current prices: Nebius raised prices “just a couple months ago” and still has “fair pipeline pressure” on supply; with 10x capacity, “not overnight, but we definitely have demand for that.” But elasticity has a ceiling — inference is a per-customer serving cost, and “there is a level where economics doesn’t work,” so Nebius optimizes customers’ total cost of ownership rather than just charging what the shortage allows.
- The strategic architecture is a four-layer stack that widens the customer base at each step: bare-metal megawatts (a dozen possible customers like Meta and Microsoft), managed multi-tenant cloud in GPU hours (hundreds to thousands of teams), managed inference in tokens via Token Factory (thousands of vertical-AI builders), and a speculative agentic layer focused on end-to-end task execution (tens of thousands of developers) — which Stebbings pegs as a direct OpenRouter competitor.
- Open source erodes frontier labs less than feared, in Chernin’s telling: builders start on OpenAI/Anthropic/Google, then at scale shift to tunable open-source models — “the most important quality of those models is not just they’re open source but they are tunable” — while the labs leapfrog to the next frontier of unsolved tasks. Stebbings’ pushback stands: trillion-dollar valuations priced to perfection while “constantly playing a game of leapfrogging… that’s a hard life to live.”
- The Revolut case is the enterprise thesis in miniature: 99% of its inference budget sat in closed OpenAI models, the shift to open source stalled on building internal evals and AI CI/CD — and once that “cold start problem” was solved, consumption grew “the same exponential trajectory” as AI-natives’ ARR. Chernin expects Revolut, Shopify, booking.com-type companies to follow.
- The biggest threat to Nebius “is not competition but consolidation”: in a world of “three, five super-models, super-empires,” Nebius is reduced to serving their physical layer. Capex this year is “2025 billion” as spoken against hyperscalers ~8x bigger; capital can’t help inside 6 months, accelerates some at 12, and “in 24 months you definitely can unlock so many things.” On Aschenbrenner’s 5.3% stake: “they give you a credit that you will execute… go back to your job and deliver.”
Deep dive
1. Not a bubble — one working use case, first percent of adoption
- Chernin refuses the bubble frame while flagging his own bias: “I probably am biased — I would probably not be in the business that we are doing if I wouldn’t believe.” His core evidence is startling in its narrowness: “we have maybe one use case that works out of so many” — coding — and it “started working like maybe a few months ago.”
- The enterprise lens does the rest: take any company in the world, even a technologically advanced one, and “you will actually see that they start using AI in the first percent of the volume, in the first percent of the use cases.” His conclusion — “even if you don’t believe what Musk says about everything in the future, space and so on, just practically from enterprise adoption, it’s just the first steps.”
2. The DeepSeek week: cheaper tokens meant more demand, not less
- The episode’s best specimen of Jevons in the wild: during the DeepSeek moment (he mentions February or March 2024/2025), Nebius stock fell 40% in one week — “the same exact week we probably had the best week in sales,” the best commercial week in company history. Customers “figured out that they can run inference in their production workloads with DeepSeek and economics will work,” and Cursor was “the first who really benefited” from tuning those models for coding.
- Chernin’s mechanism, worth keeping whole: “Every time we got the same unit of intelligence cheaper, we are not reducing the consumption but increasing the consumption — we can solve more complex tasks with the same budget, or finally economically viably solve tasks we knew were solvable but economics didn’t work.”
- This is also why open source doesn’t gut OpenAI and Anthropic: each efficiency gain pushes the frontier labs to harder unsolved tasks, and “there are so many unsolved tasks that they continue to exponentially grow.” Stebbings’ pushback — these companies are “priced to perfection at a trillion dollars” while “constantly playing a game of leapfrogging from value to value… that’s a hard life to live” — gets a pie answer, not a rebuttal: “there is enough pie for both frontier capabilities and very tuned models for specific use cases.”
3. The four-layer stack — from megawatts to GPU hours to tokens to tasks
- Nebius’s product ladder maps to units of sale and customer count. Layer one is bare metal, spoken in megawatts — “a dozen customers in the world,” Meta and Microsoft class, who “bring everything they have, deploy on your infrastructure and run.” Layer two is multi-tenant managed cloud sold in GPU hours to hundreds or thousands of research-heavy teams. Layer three is managed inference — Token Factory — sold in tokens to thousands of vertical-AI companies who “build products, they don’t do models.”
- The fourth layer is explicitly speculative: agentic workloads where “you may not even think in terms of the particular model… you want the end-to-end task to be efficiently executed” — the platform decides whether to call the smarter model or run two lighter models with a judge. Stebbings names it: “layer 4 is a direct competitor to OpenRouter.” Chernin’s differentiation claim: making agents “reliable, repeatable, and economically viable” is “a system problem,” not a model-choice problem.
- The through-line: “the higher up the stack we move, the bigger population of customers we can serve — on bare metal maybe a dozen, on managed infrastructure hundreds, on inference thousands, on agentic there will be tens of thousands.”
4. Concentration is “the main question of our business”
- Asked what revenue concentration with a Meta or Microsoft he’d accept, Chernin calls it “the main question” — not just of Nebius but of the whole product category. The stated long-term strategy is “as much diversified portfolio as possible.” With hyperscaler-class customers “you have a tiny, tiny additional value that you can provide above the physical infrastructure” — though he pushes back on the commodity label: “nothing is commodity when it comes to the real scale.”
- Stebbings sharpens it: without the full stack, “you become the capacity provider to these mega players… incredibly concentrated and very vertically focused.” Chernin’s concession-plus-hedge: “I think so… but we don’t know where the world will end up — in a world of infinite demand you may sustain even long-term selling bare metal.” The more demand-side competition, “the more picky you can be” about customers who value the platform.
- Against CoreWeave comparisons he demurs — “I don’t like to compare with others” — and offers “full-stack integration, down and up”: down into data centers, racks and servers to “squeeze more cost”; up into product to serve enterprises who “will not buy raw compute… they have data to migrate, systems to integrate. That’s the big game.”
5. Price elasticity has a floor — and TCO matters more than the sticker
- Nebius raised prices “just a couple months ago” and still has “fair pipeline pressure” on supply. But Chernin resists pure supply-demand pricing: training is a one-off cost, while “if you believe we’re moving to inference, inference is the cost of serving the customer — there is a level where economics doesn’t work.” Prices “are elastic to some extent,” and if customers’ product economics work, “they can grow and then we can grow with them.”
- His deeper claim: the market is “too much obsessed about the nominal price of capacity. You can price a GPU $3, $4, $5” — but platform quality changes real cost far more. “If you talk about inference, optimizations change the price of the tokens by an order of magnitude… people so much speak about the cost of a particular GPU, but if you do the right thing with the model you can change the price in times.”
6. Token Factory and the Revolut cold-start: enterprises compound once evals exist
- The layer-three pitch as told: you build on OpenAI, crack the use case, then “you don’t have enough margin or you want to apply your data and you cannot do it in the closed ecosystem.” You grab open-source weights off Hugging Face, a vLLM or SGLang engine — “and then it doesn’t work,” because production needs orchestration, caching, observability at hundreds-of-GPUs scale. Token Factory runs 60 open-source models with claimed inference cost cuts of up to 70% via distillation, speculative decoding and caching; with maybe “minimax 3” released and “Ultra” announced (as spoken) every few weeks, the platform absorbs the benchmarking and switching work.
- The Revolut story carries the enterprise thesis: “99% of their inference budget was in closed models, in OpenAI,” some use cases “didn’t work for them economically,” and the open-source migration stalled because they had to build the whole engine internally — “first of all they were focusing on evaluations.” His generalization: “people underestimate how important it is to build the foundation for improvements — metrics, eval mechanisms, this CI/CD process established for AI development.”
- The payoff clause: “when they solved these foundational problems, they start growing exponentially… they grow their AI consumption equal to their ARR” — the same trajectory AI-natives report. He names the next cohort: “Revolut, Shopify, booking.com — when they solve this cold-start problem, they will grow their AI adoption like crazy.”
- Stebbings’ synthesis, which Chernin accepts with a correction: yes, Nebius “takes away the plumbing” so enterprises can pipe away from closed providers — but “it’s not closed versus open… there will be a market for the smartest models of the world, the fastest models, the in-between — smart enough but cheap enough.”
7. Capital and bottlenecks are a function of the time horizon
- The capital math as spoken: “our capex program this year is 2025 billion” (the figure is not normalized) while “our competitors, hyperscalers, have like eight times bigger.” Unlimited budget changes exactly one thing: “Build faster. Data centers, and fulfill them with GPUs.”
- To Gavin Baker’s argument that permitting delays prevented a glut, Chernin answers with a timespan decomposition worth keeping: “In the next 6 months the capital cannot help — you have what you have, you need to deliver. In the next 12 months you can accelerate something… but in 24 months you definitely can unlock so many things.” Nebius builds “a portfolio of capacity” in phases — secure power and land, build data centers, fill with GPUs — “as much as possible in advance.”
- On public backlash (Stebbings cites 40 of 100 data centers now not being built when they go through planning and approvals): “This is the environment we need to work in.” Pragmatically, the portfolio is oversubscribed so one delayed site doesn’t break customer delivery; civically, he reaches for the Uber analogy — pushback comes with anything moving “too fast,” and engaging communities “is just part of your duty.” 70–75% of new mid-term capacity is being built in the US.
- On space data centers, a more optimistic view: “So many smart people are now working to make it happen… why wouldn’t I believe it will happen? If someone said even 3 years ago that we will build multi-gigawatt data centers and large interconnected compute clusters — I didn’t think like that, and we are here. It’s routine.”
8. Nvidia, sovereignty, and where the real risk-takers are
- The Nvidia power imbalance gets an engineering answer: “Nvidia is still to a big extent an engineers-driven company… if engineers in Nvidia respect your engineers, you will have the right foundation for relations” — relations that run engineer-to-engineer across hardware, software and inference layers. His summary of the whole strategy, which Stebbings threatens to title the episode: “Just do your f*ing job.”
- On European sovereign AI, his contrarian cut: the debate “was too much concentrated around megawatts and power rather than the builder layer.” Infrastructure follows demand — “we will build infrastructure if we have demand, and demand is coming from the builders” — so Europe should care about producing more companies like Lovable, Black Forest Labs, and likely Mistral.
- Asked where he’d invest across infrastructure, horizontal models, vertical models and applications: infrastructure is “a good place to be,” partly because “we kind of know what’s needed” — but “the most amazing people in this industry are those who take the risk to build end-user products… the real risk of building something people would need or not need. They are the heroes.”
9. The terminal risk is consolidation — and the shark ethos
- Finishing “the biggest threat to Nebius is not competition but…”: “consolidation in general. If you end up in a world where three, five super-models, super-companies, super-empires control the world, then companies like Nebius will be needed only to serve their needs on the physical layer. The more the world is democratized and diversified, the more we are needed.” Stebbings notes value is concentrating, not diversifying; Chernin’s answer is hope, not evidence — “so many people want to build independently… it organically creates a more diversified world. Hopefully.”
- On Leo Aschenbrenner’s disclosed position (Stebbings’ figures: 5.3% of the company, ~15% of his portfolio), the stock jumped, but the internal read is austere: “they give you a credit that you will execute… every time someone invests in us, they give us a credit and opportunity to deliver. Then go back to your job and deliver.” He credits the culture to CEO and founder (name garbled in captions): “you wake up and it’s a new day, you need to deliver, nothing is guaranteed.”
- His one self-criticism: “we could celebrate a little bit more… I don’t think we celebrate enough” — immediately swallowed by the closing image: “It’s like a shark. You’re alive when you move. So we have to move.”