Pioneers Insight Method Research Author
SF Compute: Commoditizing Compute to solve the GPU Bubble forever
Back to Episodes

SF Compute: Commoditizing Compute to solve the GPU Bubble forever

Summary

  • GPU clouds cannot inherit CPU-cloud software margins because every incremental accelerator can improve a model or generate revenue, making buyers extraordinarily price-sensitive. On a hypothetical $1 billion fleet, a 10% premium is $100 million; Conrad argues the customer could spend $50 million recreating the software and still come out ahead. “They don’t give a shit at all about your software.”

  • CoreWeave’s winning model is closer to banking or real estate than traditional cloud computing: secure long-term contracts with creditworthy customers, then finance the hardware cheaply. swyx says Microsoft and OpenAI together represented 77% of CoreWeave’s revenue, giving lenders confidence and lowering CoreWeave’s cost of capital. Calling those financial characteristics evidence of a bad cloud business misses Conrad’s point: “GPUs don’t look like that at all.”

  • The economically stable industry structure separates GPU ownership from the software layered above it. CoreWeave can prosper as “a really, really good real estate business,” while Modal can prosper without owning the underlying compute; Conrad’s warning for businesses that combine both is blunt: “If you mix these businesses together, you get shot in the head.” He suspects hyperscalers will also lose money reselling NVIDIA GPUs because their capital has higher-margin alternative uses.

  • SF Compute turns inflexible GPU commitments into a spot market where unused time can be resold, pushing healthy-cluster utilization toward 100%. A buyer can reserve thousands of H100s for an hour, continuously renew hourly reservations, or cancel a longer commitment by paying the difference between its contract and the market price. The company describes itself as “the power tool of GPU financing,” giving buyers primitives to manufacture their preferred cost, duration, and interruption profile.

  • The H100 glut reflected supply arriving in a synchronized wave, not GPU demand collapsing, and Conrad tentatively expects shortage conditions to return by winter. Infrastructure bottlenecks delayed clusters before resolving together; meanwhile, test-time inference may expand consumption far beyond a market previously dominated by a handful of consumer-model providers. He repeatedly conditions the forecast on next-generation chip rollouts: “I think I’m close to the market, but also a vague speculator.”

  • SF Compute’s endgame is a spot-price index supporting cash-settled GPU futures that let data centers hedge prices instead of forcing risk onto startups and venture funds. Conrad stresses that SF Compute is currently only a spot market; the prospective future would lower capital costs and reduce the need for giant, inflexible customer contracts. Without that hedge, he argues, compute obligations contribute to inflated pre-revenue valuations and an AI funding bubble: “Futures are the way to chill out the entire industry.”

  • Making compute tradeable requires operational standardization, not merely a financial interface. SF Compute burns clusters in with Linpack for roughly 48 hours to seven days, runs active and passive checks, secures BMC access, enforces SLAs, and is working toward automated refunds and replacement capacity. This physical complexity also underpins Conrad’s skepticism toward peer-to-peer GPU networks: absent architectural changes, “speed of light is really hard to beat.”

Deep dive

1. GPU scaling laws destroy the CPU-cloud pricing playbook

  • Conrad starts with the difference between serving web applications and training models. A company such as Gusto or Rippling buys enough CPUs to meet demand, after which another server produces “literally zero money.” Its capacity curve rises and then flatlines, so software services can sustain cloud margins.

  • Model builders face the opposite incentive. Training has diminishing returns, but Conrad says “there’s always returns”: another GPU can improve the model, while test-time inference can run longer, improve performance, serve more customers, or reduce latency. Every incremental GPU therefore has a plausible path to incremental revenue.

  • That changes procurement from “how cheaply can I obtain the CPUs I need?” to “how many GPUs can I fit inside a fixed budget?” The resulting customers buy unusually large quantities relative to their size, yet become more—not less—sensitive to unit price because every saved dollar can purchase additional useful compute.

  • Conrad’s numerical specimen is $1 billion of GPU hardware. A 10% software premium costs $100 million, enough to spend $50 million building an internal replacement and retain the balance. GPU clouds therefore cannot assume services will create the 50%-plus margins familiar from CPU infrastructure: “They will immediately switch off, if they can.”

2. CoreWeave monetized credit quality rather than software lock-in

  • Conrad’s compressed explanation of CoreWeave’s success is: “Sell locked-in long-term contracts and don’t really do much short-term at all.” The ideal customer pays upfront or is sufficiently trustworthy to pay over time, removing utilization and depreciation uncertainty before the hardware is financed.

  • swyx says Microsoft and OpenAI together supplied 77% of CoreWeave’s revenue. That concentration carries strategic questions, but it also produces financeable receivables: Microsoft is unlikely to default, while Microsoft’s backing and OpenAI’s apparent momentum improve the latter’s credit case versus a pre-seed, pre-revenue startup.

  • The hosts’ shorthand—“It’s a bank” and “It’s a real estate company”—is not an insult in Conrad’s framing. CoreWeave should not be judged as though NVIDIA GPU resale were a high-margin SaaS business. Treating it as a capital-intensive owner with contracted tenants makes its economics more intelligible and may still describe “a great business.”

  • Conrad’s risk chart places the probability of selling the GPUs on one axis and depreciation exposure on the other, with financing costs cutting across both. Long, prepaid commitments from low-risk customers occupy the best region; long contracts with possible defaulters are weaker, while short-duration, low-price sales paid over time create the highest-interest, most margin-squeezed position.

3. Short-term GPU sellers are squeezed from both sides

  • The failed cloud strategy was to charge enough early to cover later depreciation, then retain buyers through software. Conrad illustrates it with a $5 hourly selling price against roughly $1.50 of underlying GPU cost: early profits must offset the period when competition eventually pushes revenue below cost.

  • In practice, price-sensitive customers “push you down, and push you down” while hyperscalers retain the balance-sheet capacity to compress margins further. That destroyed the high-priced short-term quadrant and forced providers toward cheap, short contracts—the very structure where cash receipts fall as financing pressure rises.

  • Conrad suspects Microsoft, AWS, and Google may lose substantial money reselling NVIDIA GPUs, though he clearly presents this as intuition. They can profit from small experimenters or imitate CoreWeave with long commitments, but deploying the same capital into proprietary models, products, or competing chips could offer better margins.

4. NVIDIA benefits from a financing intermediary it does not control

  • Asked why NVIDIA did not simply create CoreWeave, Conrad gives a hedged strategic answer: NVIDIA would compete with the cloud providers that buy its chips. He also raises possible antitrust concerns about controlling too many stack layers, but marks that explanation as speculation and considers customer conflict the stronger reason.

  • Microsoft could conceivably internalize the role, but Conrad reasons that NVIDIA would not want one hyperscaler absorbing too much supply. If Microsoft, Google, Amazon, and Oracle became essentially the only buyers, their concentration would create a monopsony and give a few customers increasing power over NVIDIA’s pricing.

  • A broader group of GPU clouds keeps buyers competing for allocations. CoreWeave therefore serves NVIDIA not merely as a reseller, but as another source of demand that prevents the hardware market from collapsing into four or five dominant purchasing relationships.

  • The same logic supports Conrad’s sharpest industry prescription: separate hardware ownership from software services. CoreWeave can make money as GPU real estate, while Modal can build software without owning the fleet. For coupled operators, he invokes the Grim Reaper “knocking on the next door”: there are too few billion-dollar software add-ons to offset billion-dollar hardware exposure.

5. SF Compute emerged from an audio lab’s recurring solvency crisis

  • SF Compute originally intended to train general audio and music models, not build an exchange. The team avoided raising the roughly $50 million pre-seed round associated with initially training a model because the valuation, dilution, and burden of financing the next round looked potentially fatal—even though Conrad now thinks they might have raised it.

  • They expected to rent thousands of A100s on demand or month to month, only to be told that providers required commitments of at least a year. Conrad was initially angry; he later understood that allowing easy cancellation would leave the provider holding enormous inventory and depreciation risk.

  • Their workaround was to consume one month of a year-long contract and sublease the remaining 11. The catch was existential: about every 30 days the company owed $500,000 while holding roughly $500,000 in the bank. “If we did not sell out our cluster, we would just go bankrupt.”

  • Surviving that first year revealed a capability. The company became a GPU realtor, matching six-month buyers into year-long vendor commitments, operating bare-metal clusters, and taking a cut. Those bespoke deals produced a book of buyers and sellers that eventually became a genuine market with bids, asks, and a trading engine.

6. Resale liquidity makes otherwise irrational GPU durations possible

  • SF Compute can quote thousands of H100s for as little as an hour, even though no rational owner would deliberately finance a cluster around one-hour customers. That capacity typically exists because someone bought a longer term and is reselling an unused interval after its code fails or its needs change.

  • The same mechanism serves the customer SF Compute once was: a team that needs a large cluster for one month without underwriting the next 11. The market aggregates “little pockets of liquidity,” converting expiring or temporarily idle commitments into usable burst capacity rather than requiring each provider to accept cancellation risk itself.

  • Healthy-cluster utilization approaches 100%, Conrad says, because price falls until inventory clears. Actual utilization can be lower when InfiniBand fabric, switches, NICs, or other physical components fail, but idle functional time has a clearing price rather than remaining stranded.

  • For GPU-cloud vendors, SF Compute can fill an idle cluster while continuing to seek the preferred long-term tenant. When that tenant arrives, shorter reservations roll elsewhere. A long-term buyer can also cancel for a fee equal to the gap between its contract and the resale price, gaining flexibility without returning the underlying risk to the vendor or its lender.

7. Today’s H100 glut may precede another conditional shortage

  • Conrad rejects confident single-cause accounts of falling H100 prices. During the rapid build-out, bottlenecks appeared beyond the chip itself—InfiniBand cables, NICs, generators, or data-center capacity—so promised clusters missed their dates. Once those constraints cleared, many clusters came online together and converted shortage into glut.

  • swyx raises over-ordering and failed startups, but Conrad preserves an important distinction: excess supply does not imply declining demand. “It definitely seems like there’s more demand for GPUs than there ever was”; the market simply added supply even faster.

  • His deliberately tentative forecast is that the market could return toward shortage “by the winter,” conditioned heavily on how quickly future chips roll out. More access to market data makes him less categorical, not more: “I think I have more information than a lot of other people, and this makes me more vague of a speculator.”

  • Test-time inference is the possible demand accelerant. swyx estimates open-source AI at roughly 5% of closed AI, while Conrad recalls that the part of OpenRouter outside Anthropic, Google Gemini, OpenAI, or similar providers amounted to only about ten H100 nodes. Test-time inference could broaden consumption, though the resulting price impact still depends on new-chip supply.

8. Physical locality limits decentralization, while markets widen access

  • Conrad is “wildly skeptical” that household or broadly distributed GPUs will outperform a fully interconnected cluster using InfiniBand or its successor. Crypto may ease payments or subsidize a cold start, but token incentives eventually run out; they cannot remove latency or transform scattered cards into one tightly coupled supercomputer.

  • swyx supplies the strongest countercase: architectures could adapt through fine-grained mixture-of-experts routing, block attention, or perhaps 200 experts distributed across hardware. Conrad concedes that redesigned models might parallelize better across space, but without that algorithmic change, “your speed of light limitation” remains difficult to escape.

  • The beneficiaries of flexible centralized supply include customers traditional GPU clouds least want. Grad students graduate, grants expire, and projects churn, yet a researcher with a $100,000 grant may afford one large burst even if they cannot support a year. Conrad cites Standard Intelligence, Phind, and Schmidt Futures grantees as representative users.

  • VC-provided clusters solved a related credit problem. A startup cannot readily borrow $50 million, whereas a fund or capital partner with perhaps $1 billion in assets can; equity-for-compute was therefore credit-risk arbitrage. Conrad considers Andromeda’s supposedly $100 million cluster well-timed, but calls it a one-time edge now competed down and does not recommend blindly repeating it.

9. Hourly reservations let buyers engineer their own spot exposure

  • The apparently strange curve where one week costs more than either one day or one month reflects what is expiring now. Conrad compares immediate compute with old milk: capacity cannot be sold after its hour passes, so an idle block repeatedly drops toward a floor until somebody clears it. Curves beginning a week later look more conventional.

  • SF Compute does not describe this as preemptible compute. Each hour is firmly reserved, but an automation can repurchase the following hour and continue doing so. A workload loses capacity only when the next hour’s price exceeds its chosen limit, making interruption an explicit pricing rule rather than an opaque cloud-provider policy.

  • Conrad’s example sets a $4 maximum while buying at the market’s much lower clearing price. Historical averages for this strategy have been near $1 per hour, sometimes around $0.80; a patient job could bid only below $0.90 and preferentially run overnight. The buyer accepts price uncertainty in exchange for substantially lower average cost.

  • swyx connects those primitives to batch inference. Instead of accepting a single “half price, returned within 24 hours” product, an operator could financially engineer four- or eight-hour service levels around expected spot availability. That middle ground could suit background agents whose tasks are flexible but still need to finish within part of a workday.

10. GPU futures are intended to hedge risk, not amplify speculation

  • Conrad repeatedly corrects the terminology: SF Compute is currently an online spot market, “very explicitly not a future” or derivatives exchange. The eventual plan is to produce a credible index price, then support a cash-settled future that lets data centers lock in revenue without requiring physical delivery of GPUs.

  • His economic claim is that risk constitutes much of compute’s marginal cost. A hedge could stabilize a data center’s revenue, lower its financing rate, and reduce the need for rigid customer contracts. “Futures are the way to chill out the entire industry” sounds paradoxical only if derivatives are understood solely as speculative leverage.

  • Without a future, data centers transfer price and utilization risk through long commitments to startups; startups then transfer it to VCs through enormous funding requirements. Conrad argues this helps produce giant pre-revenue valuations that will not all return capital to LPs: “That is a bubble that’s totally gonna pop at some point.”

  • A cash-settled layer still requires trustworthy physical delivery underneath it. “These are supercomputers, not soybeans”: the spot market must operate clusters, audit performance, and prove that standardized compute actually works before its price can anchor a risk-management instrument.

11. Operational standardization is SF Compute’s real exchange infrastructure

  • Cluster acceptance begins with burn-in, commonly Linpack, for perhaps 48 hours or as long as seven days. This is an integration test of the whole assembled system, not merely an NVIDIA card check. Conrad recalls one cluster browning out under Linpack, an immediate indication that production workloads would be unsafe.

  • After burn-in, SF Compute runs realistic performance tests plus active and passive monitoring. Passive checks observe live workloads and flag hardware for later removal; active checks use idle periods for intrusive testing. Automated refunds work only partially because “there are always new genres of bullshit,” so uncaught failure modes still require manual investigation.

  • The company generally obtains BMC access, allowing remote resets and reimaging rather than acting as a passive reseller. Its engineers join customer Slack channels and may debug failures at 2 a.m.; strict vendor SLAs are passed through to buyers, with replacement capacity and automatic prorating or refunds as the company gets bigger.

  • Standard contracts use “this or better” specifications for storage and other variable components, supported by hardware whitelists. A custom UEFI shim—conceptually at the modern BIOS layer—downloads and boots a customer image, then supports Kubernetes, VMs, or other stacks. Controlling bare metal upward makes heterogeneous clusters look sufficiently alike to audit, price, and trade.

12. An anti-hype culture grew from years of founder volatility

  • SF Compute’s sparse nature-themed brand deliberately sets modest expectations: the user sees a plain page, then receives a supercomputer “for like millions of dollars cheaper” than expected. Conrad wanted the inverse of the ornate AI website whose product resolves into an ordinary SaaS application, especially when SF Compute’s own early product was rough.

  • The exception is San Francisco itself. Conrad enthusiastically markets the fog, bridge, surrounding countryside, and “massive amount of optimism,” countering depictions centered on the Tenderloin or grind culture. The calm imagery is therefore both product positioning and a sincere pitch for the city behind the company’s name.

  • That restraint follows a bruising founder path. Quirq’s mental-health retention approached zero when it worked; Room Service faced customers building distributed systems in-house; advice to “don’t die” led Conrad to try roughly 40 products across four years. He says email was not abandoned as intractable—he was burned out, and Intercom already possessed the ideal support customer base.

  • SF Compute’s closing hiring pitch mirrors its hybrid business. The company is hiring low-level Linux engineers in a mostly Rust codebase, while a fintech-oriented engineer owns ledgers, reporting, and the mandate to “not lose all the money.” Better financial plumbing ultimately means better vendor and buyer prices—and lets a researcher with $100,000 briefly compete at lab-scale compute.