Pioneers Insight Method Research Author
What Nvidia Is Getting From Groq | Sharp Tech with Ben Thompson
Back to Episodes

What Nvidia Is Getting From Groq | Sharp Tech with Ben Thompson

Summary

  • Thompson argues that technology amplifies underlying human issues rather than becoming inherently “good” or “bad.” Treating technology as something people can simply make moral or immoral assumes too much power over outcomes. Sharp separately says OpenAI should add ads to ChatGPT at some point in 2026.
  • Groq uses a compiler-first, SRAM-based architecture for extremely fast inference. Its deterministic, nonbranching calculations and precisely mapped on-die memory resemble a Formula 1 pit stop rather than the uncertainty of navigating a gas station.
  • Groq’s speed comes with a severe capacity trade-off. Its chips had 256 megabytes—not gigabytes—of memory, requiring many chips even for basic models and making large context windows and reasoning workloads difficult.
  • Inference will fragment between latency-sensitive and context-intensive workloads. Customer-service conversations and potentially real-time personalized ads favor speed; agents working independently can tolerate slower retrieval in exchange for larger context and stronger grounding.
  • Nvidia can make Groq’s niche architecture more valuable through software and supply-chain leverage. A CUDA-like abstraction could help direct workloads to the right architecture, while a newer process such as 2nm TSMC could materially improve a Groq chip.
  • The economics make Nvidia’s premium tolerable if the opportunity is large enough. Thompson cites $23 billion in free cash flow last quarter and says he does not care about the price in a market measured in tens of billions.
  • The licensing-and-hiring structure reflects an antitrust regime that may have made consequential deals easier to avoid reviewing. Groq remains independent, while Nvidia licensed its technology and, Sharp says with some uncertainty, brought over about 90% of its employees, including CEO and TPU architect Jonathan Ross. Thompson calls the result an “incredible regulatory own goal.”

Deep dive

1. Technology accelerates human problems

  • Thompson’s framing is that the things technology bothers people about are ultimately human issues that technology has exacerbated or accelerated. Treating technology itself as good or bad assumes too much power over outcomes and could produce worse results.

  • Sharp makes the separate 2026 prediction that OpenAI should incorporate ads into ChatGPT and says he looks forward to chronicling that rollout.

2. Groq makes inference operate like a Formula 1 pit stop

  • Groq predates the LLM explosion and began with a compiler rather than a chip. Its premise was to map simple, nonbranching calculations and their data locations before execution.

  • Thompson highlights the paradox that probabilistic AI rests on highly deterministic algorithms. His analogy contrasts the uncertainty of finding an available pump and maneuvering at a gas station with an F1 crew whose lines, people and equipment are precisely positioned: predefined execution is much faster.

  • Groq keeps data in on-die SRAM rather than relying on DRAM or HBM. In Thompson’s account, the fixed locations, lack of refresh requirement and reduced uncertainty make each calculation faster. He recalls Groq producing 1,000 words instantly in an earlier demo and cites a CNN demonstration of a live conversation; real-time speech is a possible use case.

3. Memory limits divide the inference market

  • The trade-off is capacity: Groq chips had 256 megabytes—megabytes, not gigabytes. Thompson says linking a “gazillion” chips could be necessary even for a basic model, while high-end models and long context windows require far more memory.

  • Sharp points to Nvidia’s Vera Rubin announcement, which included SSDs inside the larger multirack system for storing the K/V cache. Token generation repeats the same process token by token, while retaining prior tokens for a cohesive answer.

  • They expect inference to fragment rather than converge on one architecture. An agent working independently can tolerate slower fetching because the user is not in the loop; a customer-service interaction has much less tolerance for a long pause. They also discuss the possibility of generating personalized ads in real time in the years ahead.

4. Nvidia can convert Groq’s niche into a platform feature

  • The transaction is a licensing deal, not an acquisition, and Groq is to continue as an independent company with a new CEO. Sharp says—while qualifying that he has not verified it—that about 90% of the employees moved to Nvidia, including Jonathan Ross, whom he identifies as the inventor of the first TPU and Groq’s CEO.

  • Nvidia could abstract workload placement through a CUDA-like software layer: users would describe what they want to do while libraries and software handle more of the choice between Groq-style speed and memory-heavy systems built for context and data movement.

  • Nvidia also brings exceptional supply-chain leverage. Thompson contrasts Groq’s 14nm GlobalFoundries process with the possibility of a 2nm TSMC implementation, which he says could be “pretty freaking awesome”; the resulting offering would be expensive.

5. Antitrust pressure created an acquisition without an acquisition

  • Sharp calls the 4 p.m. Christmas Eve announcement one of the greatest news dumps of his life. Thompson says it appears investors and employees were paid out, including people who had joined Groq only the prior week, preserving incentives to take startup jobs.

  • Thompson dislikes the structure because aggressive acquisition reviews make normal acquisitions unattractive: “Why would you ever do a normal acquisition ever again?” He argues that treating deals as anticompetitive despite most targets being mostly failures has been destructive and may have opened a template that lets big tech avoid conventional review.

  • Groq was a niche but legitimate competitor that Nvidia could, in Sharp’s view, “effectively buy” without acquiring the company. Nvidia’s staff, supply-chain leverage and ability to improve the design could make Groq’s independent status less meaningful over time. Thompson says Nvidia’s $23 billion in free cash flow last quarter makes the premium incidental if the opportunity is as large as expected.

  • Sharp asks whether licensing deals could eventually face scrutiny, notes that this one is non-exclusive, and mentions Jensen Huang’s relationship with President Trump. Thompson warns that restricting employee movement or reviewing licenses would create further bad outcomes. He says the pattern began under the Trump administration and calls the regulatory result an “incredible regulatory own goal.”