
Vikram Sekar
Key Views & Dialogues
Why Memory Is AI’s Biggest Bottleneck with Vikram Sekar | EP 170
- 🗓️ Date:
2026-10-03| 🎙️ Show:Frictionless
Memory—not raw compute—is AI’s deepest systems bottleneck, making lower-memory architectures a potential lever across bandwidth and networking. As 144-GPU systems push power and cooling higher and copper fails at 400 gigabits, optics, DRAM-on-logic and specialized inference architectures may gain relevance, while Anthropic’s S1 points to more than $500 billion in obligated capital and compute and possible over-ordering.
View Dialogue Notes & Key Takeaways
Vikram Sekar’s central call is that memory—not raw compute—is AI’s deepest systems bottleneck. LLMs leave GPU compute unused because memory and networks cannot feed data quickly enough; a less memory-intensive model architecture could therefore ease several constraints at once: “the memory bottleneck, and therefore the memory bandwidth bottleneck, and therefore the network bottleneck.”
Consumer agents should expand inference demand, but not on a permanently dedicated-processor curve implied by the Mac mini craze. Most personal tasks can run on Sonnet-level or open-source models, while processors and memory are shared or spun up just in time. Demand still grows, but Vikram cautions that “everything will blow up because of agents, and it may not actually happen.”
AI remains a durable technology even if today’s financial commitments prove excessive. Vikram expects the technology to move “up and to the right” over ten years, yet he points to Anthropic’s S1 and says the figure they are obligated to deploy in capital and compute resources seems to exceed $500 billion. Labs could order too much memory, create a surplus, and suffer a collapse without invalidating the underlying adoption thesis.
Training remains a coherent-cluster problem, while inference is becoming a “Wild West” of architectures. Moving beyond 72-GPU, roughly 200 kW Blackwell racks toward 144 GPUs implies more than 400 kW plus harder cooling and networking; spreading 576 or 1,152 GPUs across racks then exceeds copper’s reach. Inference is less orderly: mixture-of-experts models pull different weights from different places, making data movement “complete chaos” and opening room for Cerebras, Groq, Microsoft Maia, Meta MTIA and other specialized approaches.
The near-term scaling trade is increasingly optical because copper is reaching a physical limit. Copper works inside a rack at 200 gigabits per second, but cannot span roughly 10 meters across eight racks; at 400 gigabits, Vikram says it is “completely dead.” Co-packaged optics can deliver reach with lower conversion power, but introduces reliability, replacement and mass-deployment questions for systems containing 100,000 GPUs.
The memory hierarchy should be optimized for useful work per token, not maximum tokens per second everywhere. A coding task where time-to-market matters may justify 1,000–2,000 tokens per second and expensive bandwidth, while an agent reading PDFs in the background can run many tokens slowly and cheaply. The economic question is: “What is the useful amount of work being done per token?”
HBM currently offers the best bandwidth-capacity balance, but stacking faces cost and scaling limits. With stacks already at 16 layers, Vikram questions whether 20 or 24 remains logical and highlights wafer-bonded DRAM-on-logic approaches that could offer “SRAM-like bandwidth but with DRAM-like capacity.” Yet any efficiency breakthrough may trigger Jevons’ paradox: lower resource use per task stimulates enough new usage that “we still don’t have enough computing resources.”
🔗 Original source & video: Why Memory Is AI’s Biggest Bottleneck with Vikram Sekar | EP 170