Pioneers Insight Method Research Author
No Priors Ep. 127 | With SemiAnalysis Founder and CEO Dylan Patel
Back to Episodes

No Priors Ep. 127 | With SemiAnalysis Founder and CEO Dylan Patel

Summary

  • OpenAI’s compact, reasoning-heavy open model could reset the price of intelligence below the frontier. Dylan expects America’s first globally leading open model since Llama 3.1 405B, with particular strength in code and easier deployment than Kimi. Elad Gil said OpenAI’s o1 and o3 were once priced around 4× GPT-4o per token; he expects open competition to “kneecap” margins for non-frontier APIs.
  • Inference differentiation is moving from individual kernels toward distributed systems and physical infrastructure. Sarah argues model-level optimizations will become open-source commodities; Dylan counters that DeepSeek inference can involve “160 GPUs or something like that,” over $10 million of hardware per replica, plus shared caches. They converge on single-node software commoditizing while orchestration, reliability, networking, and infrastructure remain defensible.
  • Roughly 200 neoclouds are heading toward consolidation because venture returns and infrastructure economics do not match. Operators range from poorly utilized GPU landlords to some clouds completely sold out on four-, five-, or six-year contracts; Dylan says CoreWeave often demands long-term commitments and may not quote startups attractively. Some weaker players already have cash flow below debt payments. His four paths: “go really, really big,” move into software, accept commercial-real-estate returns around 10%-15% ROE, or go bankrupt.
  • NVIDIA’s moat is a “three-headed dragon” of hardware, networking, and software execution—not merely CUDA. A specialist may need to be 2×-4× faster, then lose 20%-30% at the process node, ground on memory and networking, and finally lose its advantage when model architecture changes. Dylan therefore rates AMD GPUs, Amazon Trainium, and Google TPUs as more plausible second choices than a chip startup.
  • AI infrastructure has no single bottleneck; scarcity simply migrates as each constraint is relieved. CoWoS and HBM remain tight while optical transceivers, buildings, substations, transformers, generation, permits, and electricians join the list; Meta is even using temporary tent-like structures because buildings take too long. The investment is also stimulative: Dylan argues the US economy might barely be growing without AI capex, rising trade wages, and power assets with 15-30-year lives.
  • US export policy is a negotiation over how high in the AI value chain America can remain. Dylan’s hierarchy is to export services first, then tokens, infrastructure, and finally chips while restricting chipmaking tools—but zero GPU sales to China could trigger retaliation through rare-earth minerals. Even a Chinese chip using 3× the power may be competitive if China can provision 4× the power and rationally subsidize the downstream economic value.
  • AI companions raise questions about human connection, while competitive founder behavior can still sway judgment. Dylan would ask Mark Zuckerberg what happens when people socialize with always-on AI companions more than with one another. He also admits his view of Cognition flipped after watching its leader dominate a high-stakes poker table: “maybe he can take from the lion,” despite Dylan having done almost no product diligence.

Deep dive

1. Open weights push the commodity bar upward

  • Dylan’s forecast: OpenAI’s release would be America’s first best-in-class open model since Llama 3.1 405B, after Mistral briefly led and Chinese labs dominated for six to nine months. It may be weaker for ordinary chat, but “really good at code”; tool use could be interesting but confusing because users may not have access to the open tool-use systems on which the model was trained.

  • The rollout matters as much as the weights. Although leaked weights proved difficult to run because of unusual 4-bit details and biases, OpenAI planned to ship custom kernels and work with partners—giving inference providers an optimized stack on day one instead of telling them to build one.

  • Sarah’s premise is that model optimization eventually becomes open-source and commoditized, pushing competition down to infrastructure. Dylan’s pushback: Together, Fireworks, and other sophisticated providers currently achieve higher performance because low-level implementations vary enormously; providers using out-of-the-box software have “no market.”

  • Their synthesis separates node-level optimization from system software. DeepSeek inference may span roughly 160 GPUs and more than $10 million of hardware for one replica, with caching shared across replicas. The single node can commoditize while orchestration remains “a very ugly distributed systems problem.”

2. Cheaper reasoning threatens the non-frontier API margin pool

  • Sarah’s alternative data suggests reasoning remains constrained by cost and latency: Anthropic had eclipsed OpenAI in API revenue, but usage was primarily Claude 4 without thinking enabled, with coding the standout growth case. Customers want more reasoning yet deliberately restrain consumption.

  • Elad’s pricing example: OpenAI charged materially more per token for o1 and o3 than GPT-4o—around 4× at one stage—even though he described the underlying architecture as basically the same, apart from weights and somewhat longer average context. “They were just taking that as margin because they could.”

  • DeepSeek, Anthropic, and Google already forced prices downward; a capable American open model raises the commodity floor again. Dylan invokes the Jevons paradox: removing margin and shrinking the model could make reasoning cheap enough that total use expands rather than merely transferring revenue between vendors.

3. Neoclouds face four exits from the middle

  • SemiAnalysis keeps finding new neoclouds despite tracking roughly 200. Their economics are far from uniform: some suffer terrible utilization, while some are completely sold out on four-, five-, and six-year contracts. CoreWeave, for example, may avoid startups or quote them unattractive terms unless they commit long term.

  • Capital structure sharpens the divide. Commercial-real-estate investors can accept 10%-15% returns on equity, but those returns are unattractive to venture investors—and GPUs are short-lived assets. Some distressed operators offer startups astonishingly cheap capacity even though “their cash flow is worse than their debt payments,” making eventual failure likely.

  • Dylan’s ClusterMAX framework measures time to workload, Slurm or Kubernetes readiness, network performance, reliability, security, and availability. A few neoclouds already outperform Amazon, Google, and Microsoft, but most do not; today’s gold or platinum performance becomes table stakes within months or years.

  • The strategic exits are stark: Together and Nebius move upward into inference APIs; CoreWeave was reportedly exploring Fireworks for the same reason; Crusoe goes enormous with gigawatt sites. Otherwise, “you either have to go really, really big,” accept commercial-real-estate returns, or go bankrupt.

4. Hyperscaler margins are exposed, but cloud software still matters

  • Elad’s reason-for-being framework for neoclouds includes controlling power agreements or building at otherwise impossible scale. Meanwhile, Amazon, Google, and Microsoft retain “absurd margins” inherited from CPU, storage, and traditional-cloud economics that need not transfer to GPUs.

  • Dylan argues GPU customers largely consume open software—PyTorch, NVIDIA tooling, open models, vLLM, and SGLang—rather than proprietary cloud products that justify hyperscaler margins. Elad’s pushback is that valuable cloud software could exist, but the major clouds “have not delivered that software”; Sarah agrees.

  • The unmet product is infrastructure reliability that customers can “pay away.” Startups should not need multiple employees to run models, only to discover that an AI SaaS product occasionally fails and remains down for eight hours. Better operational abstractions remain a real software opportunity.

5. NVIDIA’s moat is a three-headed execution problem

  • Dylan calls NVIDIA a “three-headed dragon”: excellent GPU engineering, excellent networking, and software that may only be “okay” in isolation but towers over competitors. Its additional advantage is more than 20 years of ecosystem work and a mass of libraries.

  • Hyperscalers can imitate the broad architecture and win through avoided margin: Google with TPUs, Amazon with Trainium, and Meta with MTIA. NVIDIA and TPU designs are even converging in memory hierarchy and systolic-array size. A standalone startup, however, needs something genuinely different.

  • Specialization accumulates penalties. A startup may trail the latest process node by 20%-30% in cost, performance, and power, arrive about a year late to memory, and lose again on networking and supply chain. A nominal 2×-4× advantage can disappear before deployment, especially after a six-month slip.

  • Even Amazon blamed Trainium rack-integration yield for capacity arriving too slowly and an AWS revenue miss of a few percentage points. Dylan’s conditional choice is therefore AMD GPUs or Amazon Trainium as the most plausible second options, with Google TPUs strong but primarily intended for Google’s internal workloads.

6. Specialized silicon loses when models move

  • Cerebras, Groq, SambaNova, and Graphcore differed architecturally but shared a memory bet: far more on-chip memory and less off-chip bandwidth. Their chips offered roughly 10× NVIDIA’s on-chip memory while NVIDIA increased only about 30% from A100 through H100 and Blackwell. That bet ran into models becoming too large: Cerebras discovered that even its huge chip could not fit the model.

  • The opposite bet can fail too. Hardware built around very large compute units suits a dense model such as Llama 70B, but sparse expert models turn one large matrix multiplication into many small ones. DeepSeek’s small hidden dimensions can leave that specialized hardware inefficient.

  • OpenAI illustrates the information problem: its open model uses a deliberately “boring architecture,” while advantages involving long context, KV-cache usage, or other closed-model techniques remain secret. Betting on “transformers” is insufficient when the relevant workload may change before a two-year hardware cycle ends.

  • Sarah credits first-generation startups with identifying the secular workload correctly, but agrees that architectural prediction was harder. Dylan sees the same pattern at the edge: 40-50 startups lost to general Qualcomm or Intel chips adapted from smartphones and PCs. Incumbents need only move close enough to erase specialization’s edge.

7. Data-center scarcity migrates faster than builders can solve it

  • In 2023, the constraint looked simple: NVIDIA chips, then CoWoS packaging and HBM. Those bottlenecks persist, joined by optical transceivers, data-center real estate, substations, transformers, generation, grid connections, and labor. Solving one constraint merely permits perhaps 20% more output before the next appears.

  • The examples become increasingly extreme. Meta uses temporary tent-like structures because conventional buildings require too much time and labor, yet still delayed Ohio GPUs over grid issues. On-site generation then encounters what Dylan variously calls an “eight-year backlog or whatever, four-year backlog” for GE turbines.

  • Labor is equally binding. One company in a pitch claimed it had booked every contractor capable of a regional project—“we took all the people”—forcing rivals to fly workers in. Electrician wages are rising because America has too few tradespeople, while training takes longer than the current build cycle.

  • Execution consequently becomes part of model competitiveness. Dylan says Microsoft was too slow for Stargate, while OpenAI is turning to Oracle, CoreWeave, Nscale, G42, and others. He questioned what xAI had done to merit its prior funding and a valuation higher than Anthropic’s without a leading-edge model, but credits Elon and the unusually fast Colossus build: “You have to go crazy.”

8. AI capex may be holding up growth rather than causing a bust

  • Dylan flips the recession framing: without AI buildouts, the US economy might be growing only weakly or not at all. Data-center spending lifts electrician wages and funds power assets with 15-30-year lives; the capex itself feeds current economic activity while expanding long-lived capacity.

  • The White House AI Action Plan’s infrastructure ambitions therefore extend well beyond GPUs and electricity. Operating something “the size of Manhattan,” with shifting topology, novel hardware, failures, and dense networking, requires new labor capacity, imported workers, automation, or robotics—not merely more chip allocations.

  • Software has fast experimental cycles; physical infrastructure does not. Dylan’s broader call is for “hyper-competent” organizations that creatively attack every constraint, because a slow turbine order, rack-integration failure, permit dispute, or missing contractor can postpone revenue for years.

9. Export policy is a contest for the highest-value layer

  • Dylan’s Lebanon story frames the soft-power stakes: 12-year-olds whose worldview came from TikTok asked San Franciscans whether people were shot in the streets. Hollywood once spread an unintentionally positive American image; fragmented social media can now spread a very different one.

  • Models carry worldviews too—“you talk to Claude and it has a worldview.” Dylan therefore wants global users running American technology rather than Chinese models, but distinguishes that goal from any single control regime. The challenge is deciding who may operate the stack and where.

  • The prior diffusion rule favored American operators such as Microsoft or Oracle building abroad while making it difficult for smaller independent companies to build large foreign clusters. The current administration discarded that approach, yet Middle Eastern capacity still largely involves American operators or customers, including G42 capacity rented to companies such as OpenAI.

  • Dylan’s hierarchy is to sell “the highest-value, highest-margin thing”: services first, then tokens, infrastructure, and physical GPUs, while restricting sensitive chipmaking tools. But refusing China every GPU invites retaliation; a brief EDA-software restriction and rare-earth pressure illustrate the push and pull.

10. China can rationally subsidize inefficient chips

  • Sarah surfaces the implication: if forced, China may eventually build price-performance-competitive GPUs. Dylan hedges “maybe not equivalent,” but notes that a chip consuming 3× the power can still work if China provides 4× the power, even on N−2 or four- or five-year-old process technology.

  • The relevant payoff exceeds hardware margin. Model APIs can have good margins, services built on those APIs can have higher margins, and automation creates still greater economy-wide value. China could therefore rationally subsidize expensive chips while collecting revenue and gathering prompts, databases, and strategic data through its models.

  • Solar and EVs are Dylan’s analogy for patient industrial subsidy. He says the US can slow Huawei through tighter controls on equipment, memory components, and wafers, but enforcement has holes—including Huawei’s alleged access to TSMC wafers through shell companies. Finding the “massive gray line” between access and containment remains the policy problem.

11. AI companions and poker expose the human layer

  • Dylan’s question for Mark Zuckerberg is philosophical, not infrastructural: what happens when people talk to AI companions more than to other humans, especially through always-on Reality Labs devices? He asks whether the result is deeper connection or “sloppification and complete brain rot.”

  • He once considered Cognition “NGMI,” expecting general models from OpenAI, Anthropic, or xAI to overwhelm Devin; Claude Code already looked better internally. His view flipped after seeing Cognition’s Scott dominate CEOs and finance professionals at a high-stakes poker table: “maybe he can win, maybe he can take from the lion.”

  • The confession is intentionally anti-analytical: Dylan had done almost no diligence on Cognition’s code product, yet the founder’s competitive behavior changed his priors. Elad calls the Windsurf acquisition “a pretty good hand” to play, while Sarah reduces the application-investing lesson to backing “live players.”