Jonathan Ross, Founder & CEO @ Groq: NVIDIA vs Groq - The Future of Training vs Inference | E1260
Jonathan Ross, Founder & CEO @ Groq: NVIDIA vs Groq - The Future of Training vs Inference | E1260
Summary
- Groq’s headline number is being misread: “we did not raise 1.5 billion — that’s revenue,” which Jonathan Ross sizes at “about 30% of the revenue of OpenAI.” The Aramco/Saudi deal puts up the capex, Groq deploys and repays to an agreed IRR before the split flips — “we are not limited by capital anymore.” The stated goal: at least half the world’s AI inference compute by end of 2027, run on the creed that “when you are growing faster than exponential there is no amount of profit that you can make that matters.”
- Ross refuses the Nvidia-rivalry frame and inverts it: training is “a solved problem” he happily cedes — he tells his own customers to “buy every single GPU you can get your hands on” — while Groq takes low-margin (~20% upfront), high-volume inference off Nvidia’s 70-80%-margin hands. Claimed economics: more than 5x lower cost, one-third the energy per token, and “just the memory alone in the latest GPUs costs more than our fully loaded capex per chip deployed.”
- The structural wedge: Nvidia is a monopsony for HBM — only SK Hynix, Samsung and Micron make it, and it’s the hard-to-ramp part of a GPU. Groq’s LPUs skip HBM entirely and run on the same silicon process as phone chips, so “we effectively have almost no limit on how much we can scale up”: 640 chips in production at the start of 2024, 40,000 by year-end, 2M+ targeted this year — needing nearly all of its fab’s capacity.
- His biggest macro worry is a power crunch in three to four years: ~20GW of potential power against ~15GW of existing data centers worldwide, overbuilding now; when burned builders retrench and chip counts keep doubling every 18-24 months toward 120-240GW of chip-driven demand, “that power will become a hard bottleneck.” Meanwhile “fake data centers” — real-estate-focused builders with no generators or water — inflate apparent supply.
- On scaling laws: the apparent asymptote is an artifact of assuming all data is equal quality — an LLM generating and pruning its own synthetic data climbs beyond that apparent asymptote, compounding with test-time reasoning into “geometrically increasing improvement.” Compute is only a “soft bottleneck,” DeepSeek was an algorithmic tweak, and Nvidia’s 15% post-DeepSeek drop was “the popularity contest side of the market… nothing to do with the weighing machine.”
- Bubble verdict, both halves intact: “I can guarantee you that a huge amount of money will be incinerated, but I also bet that in total more money will be made than will be put in.” The novelty this cycle is a Keynesian beauty contest “gone completely amok” — multiple rivals each holding billions, so capital no longer anoints winners: “the people who have the best products are actually going to be the winners because everyone can be capitalized.”
- The positioning sermon, earned over seven years without product-market fit: “your job is not to follow the wave, your job is to get positioned for the wave.” Groq survived on “Groq bonds” — 80% of employees traded salary for equity, half down to the statutory minimum — and his map of the next waves names four era-defining companies: whoever solves hallucination, agentic sub-goals, “invention,” and decision proxying.
- The China tell isn’t Blackwell access — cloud rentals and Malaysia/Singapore “wink wink” GPU deployments may make physical access less decisive — but censorship: if China isn’t permissive of “more open, truthful models,” it’s inherently disadvantaged. And on Nvidia’s next decade he’s genuinely split: “I wouldn’t be surprised if they were 3x bigger; I also wouldn’t be surprised if they stayed around the same.”
Deep dive
1. Synthetic data extends scaling
- Ross’s correction to the scaling-laws discourse: the OpenAI curves assume all training data is equal quality, and it isn’t. His analogy — training a kid: “what’s 1+1? What’s 2×3? What’s the second derivative of the square of the hyperbolic tangent?” — is how we train models: internet dregs first, good data saved for the end. The fix is AlphaGo Zero’s move: have the LLM generate synthetic data, prune the wrong parts, retrain — “you just keep moving up,” so “the actual scaling laws don’t look like these asymptotics.”
- Why synthetic beats real: “Reddit is great, but not necessarily as high quality as talking to someone with a PhD.” A smarter model generates better data, and offline pruning makes the kept set better than the model that produced it — a ratchet, not a ceiling.
- There is a mathematical floor, though: Big-O complexity. LLMs struggle to multiply large numbers because multiplication isn’t linear — they need intermediate steps “as a mathematical requirement.” Training buys intuition (system one); reasoning is “the algorithm on top” (system two); paired, you get “polylinear… geometrically increasing improvement.” And intuition is data-hungry — three-digit multiplication needs 10x the examples of two-digit.
- On bottlenecks he rejects the singular framing: it’s compute and data and algorithms, but compute is a “soft bottleneck” — the easiest lever because it’s fungible; more of it overpowers weak data or algorithms. DeepSeek didn’t prove compute unnecessary — it was an algorithmic improvement, “this seemingly silly thing where they just wrote the answer in a box,” which made training data cheaper to generate.
2. Groq designs around HBM
- The founding bet came from two observations: at Google, every new model ended up using 10-20x as much compute on inference as training (“inference was always the critical infrastructure piece”), and AI was improving faster than Moore’s law because chip counts were also doubling every 18-24 months — 4x, not 2x. So Groq designed for unlimited chips: all model parameters live on-chip, computation flowing through 600 or 3,000 chips like an assembly line, versus a GPU running 1/100th of the line over and over.
- His Seven Powers reading of Nvidia: a monopsony — a single buyer — for HBM and the interposer (likely CoWoS). Only SK Hynix, Samsung and Micron make HBM; it’s expensive and brutally hard to ramp, and GPUs need it because regular memory would be “like drinking out of a martini straw.” The GPU’s logic die is commodity — same process as phone chips, “Nvidia actually gets it after Apple” — the memory is the only scarce part. By avoiding it: “we effectively have almost no limit on how much we can scale up.”
- The ~3x energy edge is physics: charging and discharging the long, wide wires between chip and HBM is where energy goes; on-chip memory travels short, thin wires. The corollary most people get backwards: edge computing is less energy-efficient than the data center — mopeds versus a freight train hauling a ton of coal across town.
- No switches (“we just plug our chips into our chips — our chips are the switch”), no network tuning: 51 days from contract to first tokens in production in Saudi Arabia, with Groq’s approach described as “100% predictable” versus GPU variability — his Paris-traffic-versus-trains analogy.
3. “Buy every single GPU you can get your hands on”
- Ross won’t accept the rivalry frame: “if you’re competing, you have done something seriously wrong… someone else has already solved the problem.” Nvidia does training “better than anyone else and by such a wide degree it’s a solved problem” — it just doesn’t offer fast or cheap tokens. When a demo prompted a customer to ask if they should stop buying GPUs: “no — buy every single GPU you can get your hands on… because I want your models running on us to be really good.”
- The margin math he volunteers: Nvidia at 70-80%, Groq at ~20% upfront, plus a back-end share once partners hit their IRR. More than 5x cheaper — “just the memory alone in the latest GPUs costs more than our fully loaded capex per chip deployed” — and at one-third the energy per token, the GPU’s opex alone equals Groq’s capex plus opex.
- Hence the judo framing: “you can almost say we’re one of the best things that ever happened to Nvidia” — they sell every GPU they make for high-margin training, “we’ll take the low-margin high-volume inference business off their hands.” More inference drives more training demand and vice versa; Groq has even tested selling LPUs as a “nitro boost” for customers’ existing GPU fleets.
- His ten-year Nvidia view is genuinely two-handed: “I wouldn’t be surprised if they were 3x bigger; I also wouldn’t be surprised if they stayed around the same” — the investment case assumed they’d run away with inference too, and “they just haven’t built the right thing for inference.” A fair multiple, but “the popularity contest skews everything.” He also mocks the GTC “30x faster” chart — cherry-picked curve endpoints — as classic spec-war marketing: “just tell me what the tokens per dollar is and the tokens per watt. Nothing else really matters.”
4. The $1.5B is revenue — and the business model is the second invention
- The correction he wants on record: “we did not raise 1.5 billion. That’s revenue… about 30% of the revenue of OpenAI.” Structure: the Aramco and new Saudi entity partnership puts up the capex, Groq deploys and repays from earnings to a “decent IRR,” after which the split flips — debt-like but with upside participation, and profitable for Groq upfront. “We are not limited by capital anymore… we’re limited in how much money we can make based on how much we can deploy.”
- The scale trajectory: 19,000 chips deployed in 51 days last year; selling everything possible this year makes the contract “many billions,” with tens of billions of dollars of hardware capacity available next year — “if we were talking GPU prices, we’d be talking hundreds of billions. We’re just not charging that much.”
- Against the paper claiming Groq can’t be profitable at lowest price: “we have a very positive contribution margin right now… as far as we know we’re the only ones actually making money running these open source models,” while rivals burn VC dollars fighting for share “Uber style.” Proprietary models are joining on rev-share — a Play AI voice model debuted at LEAP.
- The operating creed, repeated to the team: “we are growing faster than exponential, and when you are growing faster than exponential there is no amount of profit you can make that matters — what matters is getting a toehold.” Target: at least half the world’s AI inference compute by end-2027, holding margin steady while cutting prices — “then we get into Jevons’ paradox and life gets great.”
5. Power becomes a hard bottleneck
- Today’s demand signal is an echo: a hyperscaler tells 60 data-center builders it needs a gigawatt, “and all of a sudden there’s like 60 gigawatts of demand.” The real count as Ross tallies it: he’s aware of ~20GW people want to make available against ~15GW of data centers worldwide — more than double current capacity.
- His actual fear is the whipsaw: slight overbuild now, burned builders retrench (“I built up all this power and no one’s using it… we’re never going to do this again”) — then chip counts keep doubling every 18-24 months, turning 15GW of need into 120GW, then 240GW. “That power will become a hard bottleneck in three to four years.”
- Much of the pipeline is fake — people who think data centers are real estate. The industry joke: 100MW in three months, sign here. What’s your uptime? Where are your generators (“you know there’s a 90-month lead time on generators right now”)? Where’s the water? “Wait, data centers need water?” Amazon doesn’t fall for it; those purported projects may not be developed.
- The ecosystem runs on a duration mismatch: models amortize ~6 months, chips 3-5 years, data centers 10-15 (asks are down to 7), power plants 15-20. Yet the longest-dated asset is the least risky because it’s the most generic: “if we don’t use it for AI, we’ll use it to power all of the electric cars.” Which is why sovereign-grade, long-horizon partners like Aramco are the natural counterparties.
6. The Keynesian beauty contest has gone completely amok
- The capex tally as he recites it: Meta $65B a year, Google “70 or 75,” Microsoft 80, plus Stargate — chips and systems included. “There’s never been anything like this,” but never a case where end value was so clear: Google stayed private to hide search economics from Microsoft, and “the moment they went public — Bing.”
- The bubble verdict, both halves: “I can guarantee you that a huge amount of money will be incinerated, but I also bet that in total more money will be made than will be put in.” The charlatan phase is predictable — “AI t-shirts… AI thermal grease… next thing you know you’ll have an AI condo” — and the cure is education.
- His Keynesian beauty-contest explainer (“this will explain everything you need to know about VC”): you bet on where the money piles up, not on beauty — SoftBank’s strategy was to win by out-bidding. The unprecedented thing this cycle: “you see people raising billions of dollars who have competitors who’ve raised billions” — the contest has “gone completely amok,” so “the people who have the best products are actually going to be the winners, because everyone can be capitalized.” The cost: talent splits off to “a competitor that shouldn’t exist.”
- On value distribution he’s a power-law man — the bigger the economy, the wilder the swings toward single dominant entities — which makes the present anomalous to him: the hyperscalers’ market caps are “so closely grouped… you would expect one of them to just be killing it. I don’t understand why.”
7. Seven years without product-market fit — position for the wave
- The sermon: “your job is not to follow the wave, your job is to get positioned for the wave — and that’s the hardest thing to do, because everyone is trying to talk you into coming on shore.” Build Uber in the dial-up era and you book a ride but can’t get home; build medical or legal AI today and hallucinations kill you — but if the fix arrives and you held position, “you were perfectly positioned, just like Groq.” Almost everyone told Groq not to do LLMs: “we’re like, this is literally what we built for.”
- The near-death mechanics: close to zero cash, Groq copied WWII war bonds — “Groq bonds,” pitched at an all-hands where leadership chose vulnerability over projected strength. Instead of quitting, ~80% of employees traded salary for equity, ~50% down to the statutory minimum; when the first tranche of the $300M round closed, less cash remained than the bonds had saved. What changed for him emotionally is what PMF feels like: “the world is brighter, the birds sing… I sleep” — founders otherwise live entirely on a third kind of happiness, “future happiness.”
- Was there doubt? “There was doubt, but there was never a pause” — because the mission predates the company: AI is “the most important technology,” and unchecked it hands a few people “outsized control.” “Our goal is to preserve human agency in the age of AI. If we don’t do that, we have failed.”
- Where he splits from Dario on safety: “I’m worried about different things… people voluntarily giving up their decision-making authority.” His term is “financial diabetes” — told through his father repeatedly losing fortunes, down to talking a Chinese-food delivery man into dinner on credit outside a $20M mansion — and he expects abundance to spread it society-wide: “what happens if you can just live a life without working… how do we get people to still be making their own decisions?”
8. Four companies will define the era — plus two moonshots
- Asked to bet on one non-Groq company, he describes four unnamed types of era-defining companies, in sequence: whoever solves hallucination — unlocking medicine and law — then whoever best breaks down sub-goals for agents: “agentic comes after you solve the hallucination problem, otherwise you’ve got these long chains where you can introduce hallucinations.”
- Stages three and four: the “invent stage” — LLM writing “is terrible because it’s predictable,” and invention needs output that’s “non-obvious but obvious when you see it” — then the “proxy stage,” when you trust a model like an EA or chief of staff to make decisions for you. Money into agents now isn’t necessarily burned: Perplexity does “just fine” with hallucinations “because it’s not high risk.”
- His flagged-as-crazy prediction, hedges intact: having lost 70 pounds on Mounjaro, he thinks “if it is possible” to significantly slow or stop aging, you will have a Mounjaro moment in maybe the next 10 years — sudden, out of nowhere. Harry, not Ross, says longevity could extend by 60 years; Ross reiterates, “if possible.”
- What he’s singly most excited for: prompt engineering as a new accessibility wave — hardware was arcane knowledge, software needed machine time, but “language — you already know it, you don’t have to learn a thing.” Give 1.3-1.4 billion people in Africa a tool that builds applications by speaking: “that would be another 1.3, 1.4 billion potential entrepreneurs.”
9. Management by Big-O: 300 people, problem units, anti-founder mode
- Groq “only does things that require a sublinear number of employees” — if doubling customers means doubling headcount, the design is wrong; automate instead. The result: 300 people built their own chip, networking hardware and software, runtime, orchestration layer, compiler and cloud. His disruption template: Walmart doubles customers by doubling stores; Amazon doesn’t double websites — “don’t just say I need more people. Focus on the algorithm of your business.”
- Growth is metered in “problem units”: every tripling of anything yields the same number of problems, and management bandwidth is finite. Scaling 640 to 40,000 LPUs was four problem units; tripling headcount simultaneously would have been another — so the team stays small by design, with “talent density” a declared constant.
- His contrarian belief: “I’m going with anti-founder-mode. I believe in delegation” — micromanaging means “that person is not right for that job,” or you haven’t aligned them. Alignment is engineered: everyone carries a 25-million-tokens-per-second challenge coin, and in meetings anyone can tap it on the table — “no, no, no, that’s not the way this is going to go.”
- On the $1-2M salary wars: Groq deliberately never offers the highest number — “if we win in a bidding war, the next time someone comes along with a higher salary, that’s it. There’s no loyalty.” Equity-motivated hires are “so much easier to manage because they’re mission-oriented… they’re not there because they want the kombucha.”
10. China and Europe diverge
- Where China genuinely leads: willingness to cross lines — “they distilled the OpenAI model,” a red line most providers wouldn’t touch — and brute scale: asked to compare Stargate with China’s $128B commitment, Ross says “if they want to deploy 150 nuclear reactors, no big deal, they just do it,” so less-efficient chips can be overpowered at home. DeepSeek is “a shot in the arm for morale” and “Sputnik 2.0” — but export is harder, because the world lacks the power to run inefficient accelerators.
- The tell that decides “whether or not China has a shot” is censorship: “one of the biggest nightmares they have is free speech… can you imagine Xi Jinping saying — country, we’ve lost our advantage in AI, I need your help? Never, ever.” If China isn’t permissive of “more open, truthful models,” it’s inherently disadvantaged — and every tech founder fears becoming Jack Ma: “if your craft is AI, I’d be looking for the exit.”
- Blackwell access may be less decisive: “most of the cloud providers are happy if you swipe a credit card to rent it to you,” and the Malaysia/Singapore GPU buildout runs on “wink wink, we’re not going to rent it to China” — otherwise “that’s a lot of GPUs for that region.”
- Europe’s fix isn’t the EU’s supposedly 1,500 AI-safety hires (“I wouldn’t waste my time regulating something that doesn’t exist”) but a risk-on enclave: with Xavier Niel and Station F’s Roxanne he sketched “City F” — 10,000 people scaling toward a million, special economic dispensations, next-day job switching — because six-month waiting periods “suppress wages… and the company has to pay for the six months anyway. It makes no sense at all.” To the fairness objection about punishing incumbent insurers: “there is no right to be an incumbent, especially a slothful incumbent.”