How AI Is Reinventing Computing from Chips to Power
How AI Is Reinventing Computing from Chips to Power
Summary
- a16z is launching a new fund aimed at AI’s hardware bottleneck, and the clearest signal is founder migration: hardware pitches from top founders have jumped from roughly 3–5% of deal flow to “north of 20% or 30%.” Casado’s tell: “the founder community, which tends to be much smarter than the VC community, has identified this as a very active area.” Scope is “computer science infrastructure” — anything a model runs on, from chips and interconnect “probably to the electricity,” excluding regulated verticals.
- The anti-hype evidence is unusually concrete: Raghuram says hyperscaler capex is ~$700B this year heading toward a collective ~$1T next year; Casado says supply is “basically all booked out to 2028,” with GPUs reselling at 4x purchase price and only 5–10% of the addressable market served. Casado says “we’ve never seen prices go up on chips,” while Horowitz notes prices typically went down. At Hot Chips, the leading memory vendor said today’s demand alone would take 3 years of capacity to fill; unlike the dark-fiber era, “every GPU that’s being created is already pre-sold.”
- The core economic shift: the mythical man-month no longer governs this bottleneck — money now converts directly into capability with no natural engineering governor. Horowitz: catching a two-year lead by hiring a thousand engineers “never works… okay, now that works,” via “$3 billion and lighting up a magnificent cluster” — how a Grok or a Kimi comes “out of nowhere.” Casado expects token demand growing “close to 1,000% a year, which you cannot grow supply that fast,” while Horowitz expects compute needs to persist for decades.
- The most tradeable mental model in the episode is Casado’s per-model ASIC math: a $3–5B frontier model needs ~$10B of inference payback, so a 20% efficiency gain is worth $2B — “and you can easily build an ASIC for $2 billion.” Fixed model weights make per-model bespoke silicon plausible in a way stateful software did not, though Casado says they do not know whether the industry will go there. Hardware optimization is now “absolutely meaningful to the upside of the business.”
- The physical wall is real: rack power requirements are described as moving from roughly 5–10kW to “100 to 50 kilowatts”; AC power no longer works at that level, liquid cooling is already table stakes, only 2% of U.S. electrical engineers or electricians are certified on DC power, reinforced concrete is among the fastest-rising prices, and ~44GW of new 2028 data-center demand faces ~25GW of expected grid additions. U.S. obstacles are so severe that new companies often seek GPU capacity “in Mexico or Australia.”
- Incumbents won’t take it all: with multi-trillion-dollar silicon incumbents, “even 5% of that is a massive private company,” and NVIDIA rationally ignores silver bricks while holding gold ones. Markets expand, then fragment — the Ford/Fordlandia-to-supplier-ecosystem analogy — and desperate labs are “inking deals with companies before they actually have hardware available.”
- On agents, Casado’s frame is that the right embodiment isn’t an extension of you but “actually an employee” with its own computer and browser. Horowitz is candid he hasn’t “cracked the code”: bots “can burn a lot of tokens and spend a lot of money and get nothing productive done” — the goal is making “all our humans superhuman without wrecking the place because the bots got out of control.”
Deep dive
1. The Machine Age fund: the bottleneck moved “south of the model”
- Horowitz’s opening frame: “we have a whole new technology that’s the most important technology ever,” and every such technology demands a whole new infrastructure — “not only do we need new chips, new system software, we need new ways of doing power, we need to replace copper… it’s absolutely everything.”
- Casado’s version: infrastructure used to mean servers, storage, and network; “here it goes all the way down to the copper mines.” Models are getting better “faster and faster” and are “no longer the bottleneck” — everything “south of the model” now is.
- Casado’s founder signal: hardware used to be maybe 3–5% of top-founder deal flow (“5% is probably generous” — Horowitz says 3%), now north of 20–30%. “The founder community, which tends to be much smarter than the VC community, has identified this as a very active area.” The fund’s stated scope: “computer science infrastructure… anything a model runs on” — chips, networking, interconnect, storage, “probably to the electricity” — not regulated or verticalized industries.
- The fund adds no new GPs; the team describes it as drawing on hardware experience already in its DNA, including early checks in SpaceX, Anduril, Astranis, and Waymo.
2. Why this isn’t a hype cycle: pre-sold GPUs and rising chip prices
- Raghuram’s demand ledger: hyperscaler capex is ~$700B this year and supposedly ~$1T collectively next year — and hyperscalers see demand from everywhere: frontier labs, AI natives, enterprise, international. Application companies “are all ripping.”
- Casado adds that only 5–10% of the addressable market has been addressed. Supply is “basically all booked out to 2028,” with multi-day auctions for a few thousand GPUs and resales at four times purchase price.
- The pricing tell: Casado says “we’ve never seen prices go up on chips”; Horowitz says prices typically went down. At Hot Chips at Stanford, the leading memory vendor said today’s demand — not future demand — equals 3 years of its capacity.
- The dot-com contrast: the 1998–99 buildout was speculative dark fiber with no users to consume it and no viable video; here “every GPU that’s being created is already pre-sold.”
- Raghuram’s anecdote: a cloud-resistant large public company’s CFO found that the memory in its servers had increased so much that it could fund the entire cloud migration.
3. The mythical man-month no longer governs this bottleneck: money converts directly to capability
- Casado’s macro claim: building used to be an engineering problem with a natural governor; now “it really is a resource limitation” — “there’s nothing between the money going in and the hardware creating intelligence.”
- Horowitz on why leads no longer protect: catching a two-year lead with a thousand engineers “never works… okay, now that works” — not by hiring engineers, but “taking $3 billion and lighting up a magnificent cluster,” and suddenly “whatever, Grok can come out of nowhere… or a Kimi.”
- Why tokens keep multiplying: RL, chain of thought, and long-running agents are all inference. Raghuram says “nobody likes to use AI more than AI”; even AI writing GPU kernels to make more AI is autocatalytic. The value per unit of work is increasing, while the beneficiaries expand from developers to all knowledge workers and beyond.
- No oracle would have helped: “we’re only four years into this,” chip cycles run 3–4 years, data centers 4–5 years, and “we could not have built the capacity.”
- Casado’s horizon: the right analog is the steam engine or electricity — “probably 30–40 years of throwing compute at problems… anything with a clear reward signal,” with language and code established and computer use just beginning, before science, materials, biology, and creativity.
4. Agents as employees, not extensions of you
- Casado’s three-stage realization: AI as a feature (“it’s like a search bar”), then chat; then OpenClaw as an extension of you that “shares your keys and knows your passwords”; and finally what Clawdbot got really right: “it’s actually an employee” with its own computer and browser.
- Erik described using it over a weekend to update his credit card with several services and cancel subscriptions — “this is not coding… this is true computer use.”
- Demand sources are broadening: Erik cites a ChatGPT app with about 1 billion weekly active users and roughly 30 million developers consuming a relatively large share of compute. Raghuram says computer-use bots amount to creating half a billion knowledge workers, while embodied AI and robots are another source of demand still ahead.
- Ben’s candor on the organizational problem: these new employees “can burn a lot of tokens and spend a lot of money and get nothing productive done… they can make stuff up… they can create security problems” — and be “super duper productive.” He explicitly isn’t claiming victory: the goal is “how do we make all our humans superhuman without wrecking the place because the bots got out of control.” Ben says that after trying several approaches, Martin’s insight — “just treat them as people” — proved the most durable.
5. Rebuild the stack from first principles — possibly an ASIC per model
- Raghuram’s method: decompose the inference engine into fundamentals — what exactly is the compute (matrix multiplications), how memory should be hierarchically arranged, how chips connect on-die, across chips, and across data centers, and how much power each transmission takes — then “rebuild it from these fundamental building blocks.” That exercise is “underway in the industry right now” and is where the opportunity sits.
- Casado’s per-model ASIC math: a frontier model costs $3–5B to train; inference must pay back say 2x, or $10B. Saving 20% of that is $2B — “you can easily build an ASIC for $2 billion.” Unlike traditional software, which has state and is dynamic, model weights are fixed, making bespoke silicon per model plausible; Casado says they do not know whether the industry will actually go there.
- The model economics are unprecedented: “I don’t think in the history of the industry we’ve ever created a digital artifact with $5 billion going directly into that artifact.”
- The margin point worth keeping: software margins came free once the business worked; with AI, “the optimization in the hardware is absolutely meaningful to the upside of the business in a way that we haven’t seen in the past.”
6. Power, cooling, and concrete: the physical wall
- Rack power requirements are described as moving from roughly 5–10kW to “100 to 50 kilowatts,” which Horowitz says means “AC power doesn’t work anymore” — DC power, which needs its own cooling and is “super dangerous,” an irony given Edison electrocuted animals to attack AC. Only 2% of U.S. electrical engineers or electricians are certified on DC power; Meta now runs a free training program — “AI is going to create a lot of new electricians.”
- Many existing building designs are already obsolete: floors must bear denser racks, walls must be thicker for noise, and “once we get to Feynman, a much smaller percentage of the data centers we have today will work.” Reinforced concrete is one of the fastest-rising prices; hyperscalers are experimenting with robots to assemble and install servers.
- Casado’s scale check: a gigawatt is ~50,000 homes — his hometown, Flagstaff, Arizona (40–60k people), uses less than one — yet the host cites ~44GW of new data-center power needed by 2028 against ~25GW of expected grid additions.
- Transformers and turbines are short, permits and grid access are regulatory struggles, and new companies now seek GPU capacity “in Mexico or Australia… because it is so difficult in the United States.”
- Horowitz’s proposed fix: a standard where a data center contributes back — “power gets better, there’s no noise, there’s no water issue, and it’s adding jobs.” Some already achieve this by providing power to the state during the day and borrowing it at night, when the state’s demand is lower; their energy rates have gone down every year.
7. Why startups can still emerge alongside NVIDIA — and who builds them
- The historical precedent: Casado points to independent companies in prior epochs — Cisco and Juniper in the move to the internet, and Arista in mega-data centers — but says those changes were relatively minor compared with this one, where “everything is different.”
- The market logic: Horowitz says that continued improvement in metrics such as tokens per second per dollar, tokens per watt, tokens per rack, and power requires new innovation. Casado adds that “even 5% of that is a massive private company” when incumbents are multitrillion-dollar companies. Horowitz’s illustration is Alex Rampell’s TrialPay pitch to Facebook, where Dan Rose said: “I have so many gold bricks I can’t even pick them all up, so the last thing I’m doing is looking at a silver brick.”
- As markets expand they fragment. Casado invokes Ford’s 1913 River Rouge plant as an example of extreme vertical integration; Horowitz adds Fordlandia, Ford’s Americanized rubber-plantation city in the Amazon, which faltered when workers were made to “show up to things on time.” Likewise, “inference used to be one simple architecture, and it no longer is.”
- The founder profile differs: first rounds run hundreds of millions before product, and these must be “systems founders” who think through manufacturing and supply chain from day one. Martin says “Jensen is the Michael Jordan of this.” Raghuram notes that the memory experts “aren’t young,” while also saying all of their investments were started by two founders in their twenties who work alongside experienced people.
- Two tailwinds are that desperate labs are “inking deals with companies before they actually have hardware available,” and follow-on capital is more available.
- The name and the stakes: Martin says “artificial intelligence was the wrong word… it’s machine intelligence,” a “cache of how humans thought,” and that the name acknowledges hardware’s centrality. Horowitz’s 5–10-year win condition is “America wins in the infrastructure” with eco-friendly data centers and “an abundance of chips, memory, and power.” America, he says, is a special place where someone can come with nothing and do something profound; losing the technology lead could mean “another era and another country, and maybe they have a different set of values.”