Pioneers Insight Method Research Author
Why Big Tech Buys GPUs From CoreWeave | Corey Sanders
Back to Episodes

Why Big Tech Buys GPUs From CoreWeave | Corey Sanders

Summary

  • CoreWeave’s claimed right to exist is that AI has become too business-critical and expensive for a “best-in-suite” cloud offering to suffice. Corey Sanders compares the opening to Snowflake and Databricks during the analytics wave of maybe 10 years ago: customers may leave familiar contracts for best-in-class performance when the workload can change the business. The twist is that Microsoft and Google are themselves CoreWeave customers and, in some circumstances, partners.

  • The architecture starts with keeping the GPU—the system’s most expensive asset—continuously fed. CoreWeave optimizes object storage, Lotta Cache, multi-GPU caching, orchestration and observability around AI workloads, particularly training’s relatively limited write-out outside checkpointing, instead of accommodating every workload. Sanders’s sharpest framing: “We don’t need to design for an e-commerce website. We can design for what we’re built for, which is AI.”

  • Liquid cooling is required for some of the newest, largest GPUs and is also an efficiency thesis whose full value remains unproven. Sanders says liquid cooling can lower HVAC and air-conditioning costs, with the hope that this translates into cost savings over time. He explicitly hedges: “the verdict is still out on the full extent of the value.”

  • Corey argues that consistent or commoditized APIs do not make cloud services commodities. Operations, quality, performance and experience can continue to differentiate, while CoreWeave’s AI-specific assumptions are harder for broad public clouds to make. The bar will rise: “The level of quality, performance, capability, and experience that we deliver today will not win workloads in 2 years,” so CoreWeave must keep improving.

  • Inference could make geographic capacity more flexible because more request time is spent inside the GPU and less on the network. That could let applications spread calls across data centers, absorb bursts and improve availability without operators “always sweating” whether a particular region has GPUs. Sanders wants the platform to accept “I kind of want it here. I kind of want it there” and handle placement—but concedes this advantage might be temporary as workloads evolve.

  • Supply constraint is not one market-wide number; it depends on accelerator generation and cluster size. GB200, GB300, B200, H100 and H200 capacity are different products, as are requests for 10, 100, 1,000 or 10,000 units: “There’s not a lot of places that can run 10,000 GB200s in the world.” Customers can sometimes partition workloads, use multiple clouds or accept another generation, while on-demand and spot-like consumption should expand.

  • CoreWeave’s customer intimacy may be as important as its hardware stack, but Sanders treats the two as complementary rather than substitutes. Observability, Mission Control, CKS and Slurm on CKS simplify jobs, while customer-success teams and even the CTO work directly in customer channels—coverage Sanders says hyperscalers may reserve for their “top 9 customers or whatever it may be.” The same engagement drives products: large customers stuck trying to feed GPUs helped produce SUNK.

Deep dive

Not yet available upstream; scheduled sync will retry.