
Walter Goodwin
Key Views & Dialogues
The Future of Frontier Model Architectures with Walter Goodwin, Founder & CEO of Fractile
- 🗓️ Date:
2026-10-02| 🎙️ Show:No Priors
Fractile is betting frontier inference will be constrained by memory bandwidth and cost, not compute, shifting from SRAM to a DRAM architecture expected to be fully operational in the second half of next year. Goodwin claims 25 times more bandwidth per chip than an HBM-based design, potentially making sparse MoE models economical, while its 150-person team targets a three-to-six-month lead; foundry cycles, ramps, and three-to-five-year amortization remain constraints.
View Dialogue Notes & Key Takeaways
Fractile’s core bet is that frontier inference will be constrained by memory bandwidth and memory cost, not merely compute. Goodwin wants to combine DRAM’s capacity and economics with the “many thousands of tokens per second” associated with SRAM-based systems, enabling long-context agents and models with many trillions of parameters.
Goodwin sees much of today’s AI-chip variety as architecturally similar beneath the branding. NVIDIA and AMD GPUs, hyperscaler ASICs, and TPUs commonly rely on HBM, tensor cores, and TSMC advanced packaging. Fractile’s differentiation is attempting innovation across the entire stack rather than outsourcing physical implementation after front-end design.
Fractile abandoned an initial SRAM architecture after concluding that model size and context length would outgrow it. SRAM delivers extraordinary bandwidth but insufficient cost-effective capacity; even fast inference chips may fall back to GPUs for long-context processing. The company instead pursued “aggressively high bandwidth from the world’s cheapest memory,” DRAM, with a platform expected to be fully operational in the second half of next year.
The company’s roughly 150-person, vertically integrated team is designed to shorten the feedback loop between workload research, architecture, physical design, packaging, and foundry interaction. Goodwin argues that a durable three-to-six-month lead can decide deployments: “If you find a way to structurally secure a three- to six-month lead, you win all those implementations.”
AI may radically compress chip development, but it cannot abolish fabrication time or silicon economics. Foundry turnaround still takes three to five months, production ramps take 12-18 months, and a financially viable chip needs a three-to-five-year amortization window. Faster design therefore means maintaining more “irons in the fire,” not shipping a fundamentally new processor every few weeks.
Fractile claims 25 times more bandwidth per chip than an HBM-based design, potentially changing which model architectures are economical. Goodwin’s example is mixture-of-experts sparsity: moving from 1-to-16 toward 1-to-128 or 1-to-256 could save substantial computation, but current accelerators become bandwidth-bound. Compute has scaled roughly one million times in 20 years while memory bandwidth rose only about 40 times.
Goodwin expects frontier labs to keep buying from multiple hardware suppliers even as they develop proprietary silicon. In-house chips provide bargaining power, supply diversity, and control, but exclusive dependence is dangerous: a rival’s fivefold algorithmic efficiency breakthrough could leave a lab stranded for nine months while it redesigns hardware. Independent platforms remain valuable because they let labs compete at the model layer without betting survival on one architectural path.
🔗 Original source & video: The Future of Frontier Model Architectures with Walter Goodwin, Founder & CEO of Fractile