SemiAnalysis' Jeremie Eliahou Ontiveros on the supply/demand dynamics of AI and data centers
Summary
Ontiveros rejects “unlimited” as an investment premise, but his measured forecast is still a widening AI data-center shortage. Global IT capacity grew roughly 4 GW annually from 2019 through 2023; in 2024, NVIDIA hardware alone added about 5 GW of demand and AI accelerators collectively added 7-8 GW. SemiAnalysis expects construction to lag that demand for several more years, leaving upside across the infrastructure chain.
DeepSeek does not break the compute-spending thesis because its efficiency sits within an already extraordinary trend. Ontiveros says that, in two years, a model with GPT-3 quality should cost about 1,200 times less to run inference on. He argues Gemini 2.0 Flash is both cheaper and better than DeepSeek. DeepSeek itself remained inference-capacity constrained, illustrating why cheaper intelligence can still produce demand for much more compute and data-center capacity.
Running out of easily available text data increases the need for compute rather than ending training scale. Synthetic-data generation, reinforcement learning, stronger validation, and using huge models to improve smaller ones all consume substantial compute; Ontiveros cites GPT-4o and Claude 3.5 Sonnet as products of that broader approach. Today’s largest clusters are roughly 100,000 Hopper GPUs and 130 MW, while leading labs are planning gigawatt—and in some cases 2 GW—sites toward 2027.
NVIDIA’s underappreciated inference advantage may be scale-up networking rather than its familiar training-software moat. The GB200 NVL72 connects 72 GPUs in an all-to-all NVLink configuration at unmatched bandwidth, with 3,000-4,000 copper cables in the rack-scale system. Longer-context reasoning models require more memory and networking, leading SemiAnalysis to think NVIDIA’s moat “might actually be stronger in inference than in training.”
The AI upcycle breaks when adoption and monetization fail to support physical investment, not merely when models become more efficient. Ontiveros expects hyperscaler capex to keep beating estimates by 20-30%, but gigawatt-scale sites eventually require more than user growth. The indicators he would watch are product traffic, available indicators of OpenAI revenue and perhaps Claude revenue, and whether consumers keep paying $20 or $200 monthly. At some point, he says, the investment will be too much for the market.
The shortage makes “time to power” the most valuable infrastructure product. Schneider Electric, Eaton, Vertiv, Taiwanese liquid-cooling companies, and smaller engineering or cooling businesses can all participate; much of the equipment is more commoditized than NVIDIA’s GPUs. Behind-the-meter nuclear solutions such as Talen-style arrangements look much trickier than expected.
Bitcoin miners possess scarce powered sites, but converting them into AI campuses is a team-and-capital problem rather than a simple equipment swap. A mining facility costs roughly $0.5 million per MW, versus more than $10 million per MW for an AI data center, and customers want credible operators. Walker’s metaphor captures the optionality—miners are “clown cars that fell into a gold mine”—but Core Scientific and Applied Digital stand apart because they built data-center expertise and industry relationships early.
Miners’ power is not automatically suitable for AI. Walker flags West Texas curtailment agreements and intermittent wind and solar power as potential problems for always-on AI workloads. Ontiveros says generators, on-site batteries, and the UPS systems already used in data centers can cover outages, with battery reserves potentially extended beyond the usual 5-10 minutes.
The United States remains the default location despite cheaper overseas energy because of speed, labor, supply chains, and experience. Roughly 75-80% of hyperscalers’ large-scale self-built data centers were already American, giving the country experience with 100 MW-scale construction. Ontiveros says that no other country “currently knows better how to build large-scale data centers than the United States.”
Deep dive
1. AI has broken the old data-center growth curve
Ontiveros begins by rejecting the word “unlimited,” then quantifies why the enthusiasm exists. Worldwide data-center IT capacity added roughly 4 GW annually from 2019 through 2023; NVIDIA hardware alone generated about 5 GW of incremental demand in 2024, while AI accelerators—including custom ASICs and a small AMD contribution—generated roughly 7-8 GW.
SemiAnalysis models demand by translating every purchased GPU, CPU, or other IT equipment into the power and physical space required to operate it. Since 2024, Ontiveros says, “pretty much all” of the market has been driven by AI accelerators, making generative AI a sharp break from the industry’s previous trajectory.
Its supply-side forecast for 2024, 2025, and several subsequent years says operators are not building quickly enough. The investable implication in Ontiveros’s framing is straightforward: a growing deficit leaves room for companies supplying data-center capacity, electrical systems, cooling, and faster access to power.
2. DeepSeek sits below the existing efficiency curve, not beyond it
Walker’s central pushback is the fiber-overbuild analogy: data usage can compound for decades while investors who build at the peak still lose money. If AI developers begin optimizing power consumption as capability gains slow, today’s Manhattan-scale facilities could become stranded excess capacity.
Ontiveros answers that optimization is already proceeding at a staggering rate. He says that, in two years, a model with GPT-3 quality should cost about 1,200 times less to run inference on; DeepSeek therefore does not change the trend.
His sharper comparison is Gemini 2.0 Flash, which he describes as cheaper to serve and better than DeepSeek. The competitive achievement is real, but not evidence that frontier developers suddenly require less infrastructure than previously expected.
DeepSeek’s own constraints reinforce the point. Ontiveros says it had “zero capacity to serve inference,” turned down new-user requests, operated with a maximum batch size, and had very low interactivity. He also links its CEO’s meeting with China’s number-two political official and the following day’s announced $140 billion of state subsidies to a desire for more compute—not less.
3. Scarce real data makes training more compute-intensive
Today’s largest GPU clusters contain roughly 100,000 Hopper GPUs and consume about 130 MW of IT power. Over the following two to two-and-a-half years, Ontiveros says major labs were planning individual gigawatt-scale sites, with certain projects reaching 2 GW toward 2027.
Walker asks why clusters must grow when “most of the internet” has already entered training corpora. Ontiveros’s short answer is precisely because real data is scarce: labs must spend compute generating synthetic data, using reinforcement learning to build reasoning capabilities, and performing increasingly expensive quality checks.
Walker preserves the key failure mode: heavily synthetic corpora might recursively amplify falsehoods until a model confidently invents an Apple product launch from 1942. Ontiveros does not deny the validation problem; his answer is “scaling law again”—spend more compute to create stronger checks and higher-quality synthetic data.
Larger models also need not become economical consumer products themselves. Labs can use them to fine-tune and improve smaller models, which Ontiveros identifies as part of the path to GPT-4o and Claude 3.5 Sonnet. The frontier model’s value can therefore reside in teaching cheaper models rather than directly serving every query.
4. Untapped video offers another order-of-magnitude data pool
Before companies pay building owners to record the physical world, Ontiveros sees a much larger near-term corpus sitting unused. Models ingest YouTube transcripts and other textual derivatives, but generally do not train on the underlying video itself.
Incorporating movies and YouTube video would provide “orders of magnitude more data” than web-text sources such as Common Crawl. Walker notes that the corpus is skewed—a thousand MrBeast videos do not represent ordinary life—but Ontiveros’s point is scale: multiple years of existing video remain available before labs need schemes to pay people to generate new physical-world data.
User platforms also generate more data for their providers. Every interaction with ChatGPT, Claude, or similar platforms is stored on the provider’s servers, making broad distribution valuable as both a monetization surface and a continuing source of usage data.
5. NVIDIA’s inference moat rests on networking
Walker states the bear case in full: NVIDIA is richly valued, semiconductors are cyclical, and Microsoft, Amazon, Google, Apple, and other large customers can build competitors. They may not need state-of-the-art performance everywhere; an accelerator delivering 90% of the performance or trailing by 18 months could capture a meaningful share.
Ontiveros agrees that NVIDIA’s software advantage may matter less in inference than in training, but says the market underappreciates its overall engineering depth. Hardware and software are familiar layers; high-bandwidth networking is the third and increasingly decisive one.
His best specimen is the GB200 NVL72: 72 GPUs connected all-to-all through NVLink, at bandwidth no hyperscaler solution then matched. The rack-scale design uses “three or four thousand copper cables,” illustrating the systems-engineering burden behind what can look superficially like a collection of interchangeable chips.
Reasoning models and longer context windows require more memory; scaling memory requires a powerful scale-up network. That causal chain leads SemiAnalysis to the contrarian conclusion that NVIDIA’s moat “might actually be stronger in inference than in training,” even if its software differentiation narrows.
6. Adoption and revenue determine when the upcycle ends
Ontiveros’s cyclical stance is blunt: “If it’s an up cycle, things are going to go up, and if it’s a down cycle, things are going to go down.” Rich valuation matters less while earnings estimates keep being revised upward, which is why SemiAnalysis remained constructive on NVIDIA, power infrastructure, and other AI-exposed companies.
He expects hyperscaler capex to keep exceeding forecasts—not by 5%, but potentially 20-30%. DeepSeek does not alter that view because Google already had a model Ontiveros considered better and cheaper; its more immediate effect might be pressure on the high margins charged for state-of-the-art APIs.
What could break the cycle is deteriorating adoption. Ontiveros would monitor ChatGPT, Gemini, and Claude traffic alongside whatever indicators are available for OpenAI revenue and perhaps Claude revenue. The initial rush reflected the prospect of the next billion-user-plus platform; the physical expansion planned for the next two years ultimately needs revenue.
The analogy is Google buying pre-revenue YouTube: hyperscalers historically secure users first and construct monetization later. AI may combine increasingly capable free models such as GPT-4o Mini or Gemini 2.0 Flash with premium tools, advertising, APIs, and $20 or $200 subscriptions—but at some point, gigawatt investments require those experiments to work.
7. Time to power spreads value beyond generators
Ontiveros prefers data-center infrastructure as a capex-driven proxy for GPU deployments. Schneider Electric, Eaton, and Vertiv dominate portions of electrical and cooling equipment, while Taiwanese liquid-cooling companies and smaller local engineering or cooling-tower businesses can benefit because much of the equipment is more commoditized than NVIDIA’s GPUs.
The forecast deficit makes “time to power” especially valuable. Talen-style behind-the-meter nuclear arrangements initially looked like an answer, but Ontiveros says they proved much trickier than expected, forcing developers to consider other solutions.
He also agrees that power producers and natural gas are relevant areas to examine, but his personal focus is the infrastructure required to support GPU demand. The supply-and-demand mismatch can make solutions valuable that would not be economically viable in a normal environment.
8. Bitcoin miners own power, but AI conversion costs twenty times more
Miners historically optimized for abundant stranded electricity in West Texas, Wyoming, or North Dakota. That “power first” history created assets hyperscalers now value: a 500 MW AI project can require roughly $5 billion, while Walker’s example is that buying an existing site could bring it online in six months rather than waiting until 2029 to build it independently.
The physical mining shell is not an AI data center. Ontiveros estimates mining infrastructure at roughly $0.5 million per MW, versus more than $10 million per MW for AI, especially when the miner must outsource much of the work—about a 20-fold step-up before buying the GPUs. Customers consequently need confidence in a miner’s team, technical execution, and ability to deliver.
Walker presses the strongest objection: Core Scientific signed its landmark CoreWeave arrangement months earlier, yet few miners had secured comparable 100 MW-plus commitments despite apparently desperate demand. If the assets were such obvious “gold mines,” knowledgeable insiders should be converting faster and protecting equity rather than issuing shares aggressively.
Ontiveros offers two explanations rather than dismissing the concern. Some people in Bitcoin mining remain almost religiously committed to Bitcoin and expect it to reach $1 million; others lack the senior data-center hires and customer relationships needed for a multibillion-dollar project. “At some point, you’ve got to choose.”
9. Core Scientific and Applied Digital have the clearest execution paths
Core Scientific’s advantage was not merely power. It had an earlier relationship with CoreWeave and experience in Ethereum mining, signed a 16 MW Austin deal with CoreWeave before the flagship arrangement, hired experienced colocation personnel, and built the relationship early. That history helps explain why CoreWeave paid most of the capital expenditures for the conversion. Walker says the rumored earlier offer implied roughly a $1 billion market capitalization, while the eventual contract was worth roughly $1.8 billion in net present value.
Ontiveros calls Applied Digital the credible number two. Its 100 MW data-center capacity was already built at an approximate cost of $1 billion within a broader 600 MW North Dakota site; it had new Macquarie backing and a nonbinding letter of intent that Ontiveros thought was likely to become a roughly 400 MW deal.
IREN is his more speculative candidate because it has 1.4 GW secured in West Texas. Walker flags the area’s curtailment agreements and intermittent wind and solar power as potential problems for always-on AI workloads. Ontiveros says generators, on-site storage, and the UPS batteries already used in data centers can address outages; the usual five-to-ten-minute battery reserve could potentially be extended to 20 or 30 minutes, or whatever is appropriate.
Other examples include TeraWulf’s roughly 70 MW contract and Hut 8’s Louisiana project, which Ontiveros describes as 100 MW in 2025 and, with some uncertainty, expandable to roughly 200 MW by 2026. Ontiveros thinks Riot and Marathon likewise remain committed to mining. Raw megawatts therefore do not substitute for the teams and relationships needed to win AI contracts and access the 15-20-times-EBITDA valuations associated with colocation businesses, sometimes higher.
10. The United States keeps winning on execution speed
Walker wonders why hyperscalers do not chase nearly free natural gas in Qatar or abundant resources in places such as Brazil. Ontiveros’s answer is time to market: roughly 75-80% of hyperscalers’ large-scale self-built data centers were already in the United States, so its supply chain, labor pool, grid power, and experience support 100 MW-scale construction.
Cheap energy cannot substitute for specialized electrical, mechanical, and plumbing labor. Walker also raises the risk that a government could seize a $5 billion data-center-and-GPU investment. Ontiveros says that is a major risk, but argues that even without considering it, labor makes overseas construction difficult.
The leading AI labs, hyperscalers, and surrounding ecosystem are American, reinforcing the domestic concentration. Ontiveros’s conclusion is that no other country “currently knows better how to build large-scale data centers than the United States.”