Pioneers Insight Method Research Author
Why Memory Is AI's Biggest Bottleneck with Vikram Sekar | EP 170
Back to Episodes

Why Memory Is AI's Biggest Bottleneck with Vikram Sekar | EP 170

Summary

  • Vikram Sekar’s central call is that memory—not raw compute—is AI’s deepest systems bottleneck. LLMs leave GPU compute unused because memory and networks cannot feed data quickly enough; a less memory-intensive model architecture could therefore ease several constraints at once: “the memory bottleneck, and therefore the memory bandwidth bottleneck, and therefore the network bottleneck.”

  • Consumer agents should expand inference demand, but not on a permanently dedicated-processor curve implied by the Mac mini craze. Most personal tasks can run on Sonnet-level or open-source models, while processors and memory are shared or spun up just in time. Demand still grows, but Vikram cautions that “everything will blow up because of agents, and it may not actually happen.”

  • AI remains a durable technology even if today’s financial commitments prove excessive. Vikram expects the technology to move “up and to the right” over ten years, yet he points to Anthropic’s S1 and says the figure they are obligated to deploy in capital and compute resources seems to exceed $500 billion. Labs could order too much memory, create a surplus, and suffer a collapse without invalidating the underlying adoption thesis.

  • Training remains a coherent-cluster problem, while inference is becoming a “Wild West” of architectures. Moving beyond 72-GPU, roughly 200 kW Blackwell racks toward 144 GPUs implies more than 400 kW plus harder cooling and networking; spreading 576 or 1,152 GPUs across racks then exceeds copper’s reach. Inference is less orderly: mixture-of-experts models pull different weights from different places, making data movement “complete chaos” and opening room for Cerebras, Groq, Microsoft Maia, Meta MTIA and other specialized approaches.

  • The near-term scaling trade is increasingly optical because copper is reaching a physical limit. Copper works inside a rack at 200 gigabits per second, but cannot span roughly 10 meters across eight racks; at 400 gigabits, Vikram says it is “completely dead.” Co-packaged optics can deliver reach with lower conversion power, but introduces reliability, replacement and mass-deployment questions for systems containing 100,000 GPUs.

  • The memory hierarchy should be optimized for useful work per token, not maximum tokens per second everywhere. A coding task where time-to-market matters may justify 1,000–2,000 tokens per second and expensive bandwidth, while an agent reading PDFs in the background can run many tokens slowly and cheaply. The economic question is: “What is the useful amount of work being done per token?”

  • HBM currently offers the best bandwidth-capacity balance, but stacking faces cost and scaling limits. With stacks already at 16 layers, Vikram questions whether 20 or 24 remains logical and highlights wafer-bonded DRAM-on-logic approaches that could offer “SRAM-like bandwidth but with DRAM-like capacity.” Yet any efficiency breakthrough may trigger Jevons’ paradox: lower resource use per task stimulates enough new usage that “we still don’t have enough computing resources.”

Deep dive

1. Managed agents turn an engineering project into a consumer product

  • Logan frames the inflection as the period since Opus 4.5 in Q4 of last year, when the industry began moving from OAuth subscriptions toward API-based usage and agents began moving AI beyond chat and coding. Early OpenClaw installations demonstrated the promise—phone messages, notifications and autonomous actions—but demanded virtual machines, security decisions and enough technical confidence to expose a computer to an unpredictable agent.

  • Vikram’s own setup captured the problem: he still runs OpenClaw on a home server, yet says, “It’s not that simple.” Users bought Mac minis not always to run models locally, but to create disposable sandboxes—machines without email or bank access where an agent could “go crazy” without “blowing up my life.”

  • Cloud virtual machines removed the hardware purchase but retained the setup burden. Managed products then bundled the VM and agent: Vikram paid $200 to try Grok bot, launched agents through a polished GUI within minutes and thought, “Goodbye, $200.” That collapsing time-to-value, rather than a fundamental model breakthrough, made agents materially more accessible.

  • Instinct supplied his conversion moment. Asked to register him for the Open Compute Project Global Summit, it later returned a QR code and said it had ordered a correctly sized medium T-shirt: “If anything can do all this for me, then I’m on the team.”

2. Agent demand grows, but shared infrastructure bends the curve

  • Vikram’s caveated view—“I don’t know if I’m right”—is that personal assistants rarely need frontier intelligence. Calendar monitoring, email triage and routine administration should often work with Sonnet-level or open-source models; the giant math problem remains a separate workload.

  • Nor does each agent require a permanently assigned processor. Cloud hardware can rotate among users, memory could be taken out of service while an agent is idle, and compute could spin up just in time. Vikram’s own ChatGPT checks hourly for important emails rather than consuming dedicated compute continuously.

  • The adoption surface is nevertheless much larger because “everyone knows how to write to someone.” Parents can monitor school calendars, trips, soccer lessons and schedule changes through ordinary messages, bringing useful AI to people who never wanted a coding assistant.

  • Logan keeps the scale question in view: even shared, lower-tier inference could compound across tens or hundreds of millions of users. Vikram agrees compute keeps growing and scaling laws have not ended; his narrower claim is that agent adoption is not a straight-line translation into one frontier model and one permanently assigned processor per person.

3. Technical optimism coexists with financial unease

  • Over ten years, Vikram expects AI to move “up and to the right.” His analogy is the internet after the dot-com crash: financing and individual companies can fail while a genuinely useful foundational technology persists.

  • He keeps uncertainty around the current paradigm intact. Perhaps probabilistic LLMs are not “true intelligence” and another approach must emerge; he does not claim to know. What he can observe is that AI already lets one person manage Substack, podcasts, institutional work, meetings and rescheduling that would otherwise be too complicated.

  • Six-month forecasting is much weaker. Late last year, coding agents could have looked like the main use case; since then, concerns about chatbot half-answers and hallucinations have receded, agents have reached phones, and edge AI raises the next question: “Why do you have to go to the cloud every time?”

  • Financially, Vikram is “very nervous.” He points to Anthropic’s S1 and says the figure they are obligated to deploy in capital and compute resources seems to exceed $500 billion, while allowing that the industry might over-order memory, move into surplus and collapse. His split verdict: “Optimistic about technology,” but unwilling to dismiss excess deployment.

4. Training wants one giant GPU; inference behaves like rapids

  • Training continues pushing coherent scale-up domains from 72 GPUs toward 144, 576 and 1,152. A 72-GPU Blackwell rack is already roughly 200 kW; doubling GPU count can exceed 400 kW before accounting for the associated networking, making power delivery, cooling and physical packaging first-order constraints.

  • Spreading across multiple racks addresses the single-rack density problem but lengthens connections from feet to meters. Covering roughly eight racks requires close to 10 meters of reach, which copper cannot deliver at 200 gigabits per second. The aspiration to make many GPUs behave like “one giant GPU” consequently turns interconnect into the limiting system.

  • Vikram contrasts training’s data movement to ocean waves: operations arrive in sets, all-reduce returns the wave, and the next iteration begins. Mixture-of-experts inference looks instead like rapids—each request activates only tens of billions of parameters from a much larger model, and successive requests may need weights stored at opposite ends of the room.

  • That randomness explains why inference supports more architectural experimentation. Google has separated training and inference TPU designs since version 4, with version 8 seemingly releasing both together; Cerebras, Groq, Microsoft Maia, Meta MTIA and OpenAI’s “Jalapeño” attack the workload differently because “it’s the Wild West.”

5. Copper’s fading signal forces AI toward optics

  • Distance is the underlying physics. HBM beside a GPU can deliver about 22 terabytes per second because bits travel millimeters; across a rack they travel feet, losing signal energy and forcing lower speeds, pre-emphasis, error correction and DSPs that consume both time and power.

  • At 200 gigabits per second, copper still works within a Rubin rack. Vikram’s categorical call is that at 400 gigabits, a signal five meters away becomes nearly indistinguishable: “Copper cable has reached its limit.” It also cannot connect the neighboring racks needed to construct larger scale-up domains.

  • The power-density comparison shows why simply enlarging racks fails. Cloud-era racks consumed roughly 20 kW; an AI rack can consume ten times that, and doubling again creates a nominal 400 kW rack. The practical interim design is eight adjacent racks acting as one, which makes optical reach unavoidable.

  • Vikram presents traditional pluggable optical modules as an unattractive energy trade-off that helps explain why NVIDIA did not adopt optics earlier. Co-packaged optics moves electrical-to-optical conversion beside the GPU or switch, addressing power and reach, but raises a harsher operational question: when a tightly integrated optical component fails, how is it replaced across a 100,000-GPU deployment?

6. Optical scarcity creates a “wide and slow” design contest

  • Today’s 1.6-terabit optical connection commonly combines eight lanes at 200 gigabits per second. Those high-rate lasers use indium phosphide—the “Ferrari technology” capable of connecting data centers or continents—despite the immediate job being communication with a neighboring rack.

  • Supply is the catch. Indium-phosphide lasers come from small three- or four-inch wafers that Vikram likens to tiny pancakes; an industry built for long-haul, metro and submarine links “was never ready for this level of demand.”

  • The alternative is “wide and slow”: use many more lower-rate lanes. Gallium-arsenide VCSELs may operate around 20–50 gigabits per second; at 50 gigabits, 32 lanes can recreate a 1.6-terabit aggregate with a more affordable supply chain. MicroLEDs offer another light source but, at roughly 3–5 gigabits per second, would require hundreds of lanes.

  • Vikram gives no false precision on the winner: “I don’t know the correct answer.” The active debate is whether fewer scarce Ferrari lasers or many cheaper lanes best balance bandwidth, manufacturing, energy and reliability. Logan then broadens the discussion to power—from generation through inverters, UPS systems, storage and capacitors—as another enormous constraint.

7. Useful work, not peak speed, should determine the memory tier

  • The historical hierarchy—SRAM cache, DRAM, SSD, hard disk and tape—has fragmented. AI now spans on-chip SRAM, HBM and stacked DRAM, performance-tuned single-level NAND, capacity-oriented QLC NAND, intermediate combinations, disks and almost infinitely capacious but extremely slow archival tape.

  • Vikram’s governing question is not which tier wins, but “What is the useful amount of work being done per token?” The industry now boasts 1,000 or 2,000 tokens per second as it once boasted FLOPS, yet a background agent processing PDFs has no reason to buy maximum immediacy.

  • Enterprise coding can justify premium speed because being first to market repays the cost. Background research can run more tokens slowly and cheaply. Mapping each workload onto an appropriate memory tier should reduce cost, improve margins or pass savings to users—and thereby increase total usage.

  • Logan proposes larger context windows, perhaps from 1 million toward 10 million or an eventual 100 million as Dario reportedly suggested. Vikram’s counter is that not all context is useful. Logan’s fitness example makes the point: an agent answering a steps question can consult calorie records and fitness-tracker data without retaining work conversations about data-center optics.

8. DRAM innovation matters now, but new equations could reset the stack

  • Vikram’s preferred layer is DRAM because HBM currently offers the best balance of bandwidth and capacity: “At the moment, you can’t do without HBM.” Yet stacks are already at 16 layers, and he questions whether proceeding mechanically to 20 and 24 remains economical or technically sensible.

  • Bonding memory across the face of logic could unlock far more connections than routing through chip edges. He cites Cerebras bonding a full DRAM wafer onto its wafer-scale engine, expected work from Groq, D-Matrix’s Raptor engine and Qualcomm designs with two or four DRAM layers—the promise being “SRAM-like bandwidth but with DRAM-like capacity.”

  • The private-company landscape reflects inference’s unsettled architecture: Vikram highlights D-Matrix, SambaNova’s TCO-oriented customer proposition, Etched’s intriguing low-voltage-pin architecture, plus MatX, Fractile and Nubis’s inter-chip nanolasers. His enthusiasm is exploratory, not a declaration that one design has won.

  • The largest five-year risk sits above every hardware supplier. Because current hardware implements a particular family of LLM equations, a research breakthrough delivering comparable performance with much less memory could ease bandwidth and networking simultaneously. But Jevons’ paradox remains: every efficiency improvement may produce more use and the same refrain—“Sorry, we still don’t have enough computing resources.”