Pioneers Insight Method Research Author
Back to Pioneers
Raghu Raghuram
Founders 4 Curated Dialogues

Raghu Raghuram

VMware · Former CEO

Frontier Insights

Core Thesis: AI models and infrastructure are transitioning into sovereign geopolitical assets where compute demand will persistently outstrip physical capacity, fueled by agentic, continuous inference workloads.

Strategic Decisions: Win through vertical co-design across silicon, network, and dynamic model routing. Leverage incumbent distribution to accelerate deployment cycles while targeting anchor accounts and sovereign ecosystems rather than building commoditized, thin-wrapper applications.

Risks & Warnings: Hard physical constraints—power, land, and networking bottlenecks—will throttle growth for 3–5 years. Past tech cycles prove dominant abstractions inevitably decay; over-indexing on proprietary silos without interoperability risks displacement by open, developer-centric alternatives.

Key Views & Dialogues

How AI Is Reinventing Computing from Chips to Power

  • 🗓️ Date2026-08-28 | 🎙️ Show:The a16z Show

a16z’s new AI infrastructure fund captures a founder migration into hardware, with top-founder hardware pitches rising from roughly 3–5% to “north of 20% or 30%.” Hyperscaler capex, booked-out GPUs and resale premiums support opportunities across chips, power and cooling, while grid shortages, regulation and uncontrolled agent spending remain constraints.

View Dialogue Notes & Key Takeaways
  • a16z is launching a new fund aimed at AI’s hardware bottleneck, and the clearest signal is founder migration: hardware pitches from top founders have jumped from roughly 3–5% of deal flow to “north of 20% or 30%.” Casado’s tell: “the founder community, which tends to be much smarter than the VC community, has identified this as a very active area.” Scope is “computer science infrastructure” — anything a model runs on, from chips and interconnect “probably to the electricity,” excluding regulated verticals.

  • The anti-hype evidence is unusually concrete: Raghuram says hyperscaler capex is ~$700B this year heading toward a collective ~$1T next year; Casado says supply is “basically all booked out to 2028,” with GPUs reselling at 4x purchase price and only 5–10% of the addressable market served. Casado says “we’ve never seen prices go up on chips,” while Horowitz notes prices typically went down. At Hot Chips, the leading memory vendor said today’s demand alone would take 3 years of capacity to fill; unlike the dark-fiber era, “every GPU that’s being created is already pre-sold.”

  • The core economic shift: the mythical man-month no longer governs this bottleneck — money now converts directly into capability with no natural engineering governor. Horowitz: catching a two-year lead by hiring a thousand engineers “never works… okay, now that works,” via “$3 billion and lighting up a magnificent cluster” — how a Grok or a Kimi comes “out of nowhere.” Casado expects token demand growing “close to 1,000% a year, which you cannot grow supply that fast,” while Horowitz expects compute needs to persist for decades.

  • The most tradeable mental model in the episode is Casado’s per-model ASIC math: a $3–5B frontier model needs ~$10B of inference payback, so a 20% efficiency gain is worth $2B — “and you can easily build an ASIC for $2 billion.” Fixed model weights make per-model bespoke silicon plausible in a way stateful software did not, though Casado says they do not know whether the industry will go there. Hardware optimization is now “absolutely meaningful to the upside of the business.”

  • The physical wall is real: rack power requirements are described as moving from roughly 5–10kW to “100 to 50 kilowatts”; AC power no longer works at that level, liquid cooling is already table stakes, only 2% of U.S. electrical engineers or electricians are certified on DC power, reinforced concrete is among the fastest-rising prices, and ~44GW of new 2028 data-center demand faces ~25GW of expected grid additions. U.S. obstacles are so severe that new companies often seek GPU capacity “in Mexico or Australia.”

  • Incumbents won’t take it all: with multi-trillion-dollar silicon incumbents, “even 5% of that is a massive private company,” and NVIDIA rationally ignores silver bricks while holding gold ones. Markets expand, then fragment — the Ford/Fordlandia-to-supplier-ecosystem analogy — and desperate labs are “inking deals with companies before they actually have hardware available.”

  • On agents, Casado’s frame is that the right embodiment isn’t an extension of you but “actually an employee” with its own computer and browser. Horowitz is candid he hasn’t “cracked the code”: bots “can burn a lot of tokens and spend a lot of money and get nothing productive done” — the goal is making “all our humans superhuman without wrecking the place because the bots got out of control.”

  • 🔗 Original source & video: How AI Is Reinventing Computing from Chips to Power

Listen to full conversation →


Ben Horowitz on the Global Race for Tech, Power, and Influence

  • 🗓️ Date2026-07-03 | 🎙️ Show:The a16z Show

AI models are becoming geopolitical infrastructure because their defaults will mediate products while encoding contested histories, ethics, and values, making model choice a question of sovereignty as well as performance. Private AI, autonomy, and cyber vendors now shape allied deterrence, while global APIs expose startups to demand before local distribution exists; anchor customers and trusted relationships can justify $5 million or $10 million market-entry investments.

View Dialogue Notes & Key Takeaways
  • AI models are becoming geopolitical infrastructure because they will mediate nearly every product while encoding contested histories, ethics, and values. Horowitz’s warning that “the models are not objective. They have opinions” makes model choice a sovereignty decision, not just a benchmark or cost comparison: the defaults inside cars, education, and household systems will project somebody’s worldview.

  • Deterrence increasingly rewards innovation velocity, not only military scale, pulling private AI, autonomy, and cyber vendors into the core of allied security. Neuberger points to the Strait of Hormuz and Red Sea: adversaries can field cheap, software-built systems, while “those technologies today are not being built by governments.” Government-private-sector access and allied interoperability therefore become strategic assets.

  • AI and APIs globalize product demand before startups have the organizational capacity to serve it, creating a distribution bottleneck for venture-backed companies. Raghuram contrasts the old threshold of “a few hundred million dollars in revenue” with today’s earlier international pull, but stresses that local relationships and market structure still matter. In top-heavy economies, five or 10 companies plus government can matter most, so access may outperform a premature full-country rollout.

  • An anchor customer can reverse the economics of market entry: a $5 million or $10 million opportunity can justify the same scale of upfront country investment. Horowitz argues this is a16z’s edge in allied, AI-forward, relationship-heavy markets, where government, business, investors, and adoption are intertwined. The target map includes Japan, Korea, the Middle East, Mexico, and Canada, while Raghuram notes Japan, Korea, and Taiwan contain 15–20% of the Forbes Global 2000.

  • Cybersecurity is the clearest dual-use AI opportunity and the hardest policy trap: finding a vulnerability enables both patching and exploitation. Neuberger believes the models may help defense “far more,” because defenders cover a broad expanse while attackers need one opening; AI can “jiggle every doorknob continuously and at scale.” Yet Horowitz warns that restricting vulnerability discovery could also prevent defenders from auditing and patching their own code.

  • Silicon Valley is not software that can simply be copied online; its moat combines technical talent, entrepreneurship-friendly rules, and a culture that grants status to risk-taking. Horowitz calls the internet-as-distributed-Valley thesis “happy talk” and warns the culture is “so easy to destroy.” The investable corollary is that tax, property, hiring, and social incentives can expand—or abruptly shrink—the founder and growth-capital pipeline.

  • 🔗 Original source & video: Ben Horowitz on the Global Race for Tech, Power, and Influence

Listen to full conversation →


Building the Real-World Infrastructure for AI, with Google, Cisco & a16z

  • 🗓️ Date2025-10-29 | 🎙️ Show:The a16z Show

AI infrastructure demand is outpacing deployment: Google’s seven- and eight-year-old TPUs remain at 100% utilization, while power, permitting, land, and supply chains may constrain trillions in committed spending for 3–5 years. Specialized silicon, distributed networking, and inference-native systems could deliver 10–100x efficiency gains, but bursty workloads and the need for full-stack co-design make utilization, architecture, and durable product differentiation central risks.

View Dialogue Notes & Key Takeaways
  • Amin Vahdat sees the AI buildout as “100x what the internet was,” while Jeetu Patel calls it the internet, space race, and Manhattan Project “all put into one.” Patel says the buildout is being grossly underestimated because geopolitical, economic, national-security, and speed imperatives are arriving together.

  • Google has seven TPU generations in production, and its seven- and eight-year-old TPUs still run at 100% utilization. Vahdat says power, land transformation, permitting, and supply-chain delivery are the constraints; even if trillions are committed, buyers may be unable to “cash all those checks” for 3–5 years.

  • Power scarcity is already dictating where data centers get built and how they connect. Patel says enterprises remain early in the required re-racking and power-density transition, while Cisco has launched “scale-across” systems intended to let facilities act as one logical data center even when separated by a distance rendered in the transcript as “8 900 km.”

  • The next computing stack will be co-designed from silicon through software and will be “unrecognizable” within five years. Vahdat calls this a “golden age of specialization”: for certain computations, TPUs deliver 10–100x the efficiency per watt of CPUs, but even the best teams need 2½ years to move a specialized architecture from concept into production.

  • Networking is becoming both a primary bottleneck and a force multiplier for scarce compute power. Workloads can swing by tens or hundreds of megawatts between computation and communication, yet the most expensive network capacity may be needed only 5% of the time—an unresolved architecture and utilization problem.

  • Inference efficiency is improving by 10x and 100x, but users consume the gain by demanding smarter models and longer autonomous runs. “Intelligence per dollar” rises while total cost can still increase; prefill, decode, reinforcement learning, and inference-native infrastructure therefore create different hardware, memory, latency, and networking tradeoffs.

  • AI-assisted engineering is moving from demos into migrations that previously looked economically impossible. Google once estimated a Bigtable-to-Spanner migration at “seven staff millennia” and abandoned it; Vahdat says AI has since assisted instruction-set migration across Google’s codebase. Patel hopes Cisco’s 25,000 engineers can reach 2–3x productivity within a year, while warning that tools dismissed today must be retested within four weeks.

  • Patel’s startup warning is that “thin wrappers” around another company’s model will have short-lived durability. He favors models tightly coupled to product feedback plus dynamic routing between a product’s models and foundation models; Vahdat expects agents and productive image-and-video systems—not merely better text chat—to become transformative over the next 12 months.

  • 🔗 Original source & video: Building the Real-World Infrastructure for AI, with Google, Cisco & a16z

Listen to full conversation →


From the Dot-Com Crash to the AI Era: How Builders Survive Waves of Disruption

  • 🗓️ Date2025-08-06 | 🎙️ Show:The a16z Show

VMware’s cycle shows how a software abstraction can disrupt an incumbent before AWS retains the virtual machine and captures developers, a constituency VMware “had no idea how to work with.” Cisco’s reset targets market in nine months, $1 billion in three to four years, and 8–10 repeatable winners by combining protected two-pizza teams with scaled distribution. AI could expand infrastructure demand 100–1,000x as agents create sustained inference workloads, but vertical integration must remain open enough to include competitors such as Microsoft Teams.

View Dialogue Notes & Key Takeaways
  • Raghu Raghuram’s VMware history is a clean incumbent-cycle model: one decade disrupting, then one decade being disrupted. VMware won through a software virtual-machine abstraction and software economics; AWS retained the virtual machine but unlocked infrastructure for developers, a constituency VMware “had no idea how to work with.” His warning: “When you’re talking only to your best customers, by definition that’s not where the disruption is coming from.”

  • Jeetu Patel’s Cisco reset pairs startup impatience with incumbent distribution. The operating target is “the world’s largest startup”: move a product from zero to market in nine months, build it to $1 billion within three to four years, and repeat that 8–10 times. The scarce capability is not experimentation but doubling down—and ensuring each winner can ride Cisco’s existing route to market.

  • Displacing an entrenched vendor requires a sharply bounded customer and a genuinely asymmetric product. VMware’s vSAN initially struggled when sold to storage buyers, then worked when positioned as an expansion of the compute buyer’s remit; Raghu demands “10x better,” while Martin warns that most purported 10x gains are really 15%. NSX shows the other path: it targeted a different networking user and was pursued inorganically, with a value proposition based on capabilities physical networking could not provide. In brownfield markets, the practical sequence is “coexist first before you can displace.”

  • Organizational transformation starts with product, truth and a story that cannot be delegated. Jeetu protects two-pizza teams from corporate “antibodies,” leaves titles outside design reviews and uses binary language when a unit is failing. At 95,000 employees, narrative becomes operating infrastructure: “The story is the strategy.”

  • AI could expand infrastructure by orders of magnitude because agents turn inference into sustained demand. Jeetu’s example: 1,000 employees plus 10,000 agents creates a network load analogous to serving 11,000 employees; a delayed packet leaves an expensive GPU idle, “like burning money.” Raghu calls AI a 10x market expansion, while Jeetu argues infrastructure scale and market size could be 100–1,000x.

  • Vertical integration may matter in AI, but closed ecosystems risk self-exclusion. Cisco links its silicon, networking, security, data and observability assets while integrating competitors such as Microsoft Teams, which Jeetu says added hundreds of millions in revenue. His rule: refusing to integrate with a vendor above 20% market share excludes you from the market; Raghu nevertheless sees open models and separate inference platforms as evidence that horizontalization remains unresolved.

  • For builders, the prescription is to run toward the AI disruption while retaining commercial discipline. Jeetu ranks timing first, followed by market, team, product, brand and scaled distribution; product itself must create love, retention and commercial relevance. The closing exchange rejects imminent human irrelevance: Raghu points to cancer and paying taxes, Martin to a full board presentation and careful customer email, and Jeetu to the DMV. Jeetu’s advice is to “run toward the fire.”

  • 🔗 Original source & video: From the Dot-Com Crash to the AI Era: How Builders Survive Waves of Disruption

Listen to full conversation →