The Inference Revolution: Groq, Nvidia and the Future of AI
The Inference Revolution: Groq, Nvidia and the Future of AI
Summary
- At the Irish Open conference, the Groq founder/CEO and Google TPU inventor — introduced as Nvidia’s chief software architect, having “been in the news a little bit lately” — sat down with hedge-fund manager John. His core framework: there is no permanent bottleneck in AI — “every time a bottleneck gets big enough, people solve it,” so components can charge a premium while they’re a problem but not a huge one; as they become bigger problems, people start solving them.
- The direct application to today’s tightest trade: memory’s pricing power is self-limiting. Memory was “the most commoditized segment of the semiconductor supply chain” and is discussed through a Giffen-good analogy rather than a luxury good — you can raise the price of rice only until people switch to corn. Deep Seek’s V4 release was cited as an example of algorithmic efficiency, reportedly compressing what was likely KV cache by 90%: “it’s the tall poppy. As soon as it gets too tall, it gets chopped down.”
- He rejects the Jevons-paradox defense of memory scarcity on opportunity-cost grounds: engineers working on that problem represented an opportunity cost — “if they had enough memory, they would have worked on something else.” His prescription is blunt: “start building more memory fabs.”
- The pair’s long-standing disagreement is the thesis-relevant one: John argues intelligence has diminishing returns above PhD level and that open-source models are roughly 6 months behind closed-weight model labs, so the open-weight landscape catches up to the closed-weight landscape if his diminishing-returns argument holds. Jonathan’s rebuttal: “there’s no way to satiate the appetite for intelligence” — unsolved problems (cancer, aging), insufficient compute, and human competition mean even imperceptible model gaps show up “in my returns.”
- Agentic AI hardens the moat for the smartest model: “AI likes to use AI” and will recognize and use the smarter model even when humans can’t tell the difference. Supporting anecdote: LLM-screening recruiters prefer resumes written by their own model — so write “one resume with Claude Opus 47 and one with ChatGPT.”
- His case against a capability plateau underwrites the build-out: models now generate, prune, and retrain on their own data, “improving at a pretty linear rate,” and AI’s intuition already exceeds ours — Waymo’s daily data, he said, was starting to approach a human lifetime of driving data, though he didn’t know whether it had reached that amount. Bonus definition worth keeping: intelligence is a stationary capability—the ability to make a prediction or influence an outcome; sentience is “your rate of improvement in your intelligence” — “a property of a civilization,” and an accelerating feedback loop in society.
Deep dive
1. Two moonshots in one: the stack is bigger than any single bottleneck
- Jonathan’s opener: We landed people on the moon about 60 years ago, while AI is “cutting edge even for today” — the algorithm has existed since the 1970s, and there is finally enough compute. Investors fixate on one layer, but making AI work takes “the chips, the packaging, the systems, the networking, the data centers… power, all the stuff.”
- On top of that, training and inference are “two completely different problems” — like getting to the moon versus landing on it. John’s framing to open: portability across architectures is diminishing, and porting is “no longer an engineering inconvenience” but an economic handicap.
2. Bottlenecks that get too big get solved — memory is the tall poppy
- The core dynamic: “Every time a bottleneck gets big enough, people solve it.” A limited-but-tolerable component can charge a lot; become too big a problem, and people start solving it.
- Memory — today’s biggest AI bottleneck, formerly “the most commoditized segment of the semiconductor supply chain” — is discussed through a Giffen-good analogy rather than a Veblen good: like rice, raising the price makes buyers spend more on it, until “someone goes, let’s stop eating rice and let’s eat corn.”
- John raised Deep Seek’s V4 release, saying it compressed what was likely KV cache by 90%, and the Jevons-paradox argument around that. Jonathan’s answer is opportunity cost: “If they had enough memory, they would have worked on something else… It’s the tall poppy. As soon as it gets too tall, it gets chopped down.” His fix: “start building more memory fabs.”
3. The long-standing disagreement: do returns to intelligence diminish?
- John’s side: past PhD level, humans “can’t really understand the differences” between models — and with open-source models roughly 6 months behind closed-weight model labs, the open-weight landscape catches up to the closed-weight landscape if returns diminish.
- Jonathan’s rebuttal — “I might surprise you a little”: “there’s no way to satiate the appetite for intelligence.” As long as cancer isn’t cured, people die of old age, and there isn’t enough compute to run some AI models, there isn’t enough intelligence; and competition — “everyone in this room probably could retire… but you keep investing” — means indistinguishable models still separate in your returns.
- The agentic kicker: “AI likes to use AI,” and it will recognize and use the smarter model even when humans can’t. The resume study as told: LLMs prefer resumes generated by themselves, and recruiters now screen with LLMs — so build “one resume with Claude Opus 47 and one with ChatGPT.”
4. Why capability won’t saturate: intuition plus the synthetic-data flywheel
- Via Kahneman’s thinking fast/slow: AI is the intuitive kind, and “it’s actually better at being intuitive than we are” because of data scale. Waymo’s daily data, he said, was starting to approach a human lifetime of driving data, though he didn’t know whether it had reached that amount — you’ve seen the I-beam falling off the freight truck “two to three times and you know exactly what to do.”
- Training now bootstraps: models generate data, prune it because they “can tell what good is,” retrain, and move up — improving “at a pretty linear rate,” so “there’s no reason to believe that they’re not going to get any smarter,” even if we can’t perceive the gap.
5. Sentience is a property of a civilization, not a model
- His definitional hobby: intelligence is your stationary ability to make a prediction or influence an outcome; sentience is “your rate of improvement in your intelligence” — a spectrum, not a binary. The world’s best Go player asymptoted “because he didn’t have better players to play against.”
- Language transfers information across civilization: “intelligence is a property of an organism or an individual, sentience is a property of a civilization.” AI is contributing to that accelerating feedback loop in society — “I would expect our kids to get much smarter than we ever were.”