Pioneers Insight Method Research Author
Vol.204 Industry Watch 38 | Nvidia’s $20B Groq Deal: Where Is the Opportunity for China’s AI Inference Chips?
Back to Episodes

Vol.204 Industry Watch 38 | Nvidia’s $20B Groq Deal: Where Is the Opportunity for China’s AI Inference Chips?

Summary

  • The episode reads Nvidia’s roughly $20B deal for Groq as a major bet on the “real-time inference” phase, not simply a talent grab or an attempt to kill a rival. 杨斌 says the sum, based on the episode’s estimate, equals roughly one-quarter to one-third of Nvidia’s annual cash flow; Nvidia generated $56B in cash flow in the first 3 quarters, making this a major draw on its reserves. The deal fills the real-time inference gap alongside the Grace CPU, Rubin GPU and Mellanox DPU, while fitting Nvidia’s second growth curve into embodied-AI applications such as autonomous driving and robotics since 2025.
  • 杨斌’s architectural view is unequivocal: the LPU is “the most suitable architecture for inference, bar none.” The episode cites Groq’s commercial progress as supporting evidence: roughly $90M in revenue in 2024, an expected $500M in 2025, customers including Meta and compute centers in Saudi Arabia and Norway, and around 2.5M registered cloud users. Nvidia’s willingness to absorb its IP and core team is seen as a mature strategic decision.
  • What makes the LPU commercially valuable is not just low absolute latency, but low latency that stays low over time. 杨斌 cites Groq’s comparison presented at Hot Chips 2024: against Nvidia’s 4nm H100, its 14nm chip delivered roughly one-sixth the latency, one-third the power consumption and one-quarter the cost, with Groq claiming 10x higher overall energy efficiency. “Fast one moment, slow the next” hurts the chatbot experience; in autonomous driving, it can determine whether the use case works at all.
  • Groq’s breakout after 9 years was not driven by a sudden improvement in technology, but by models, applications and commercial economics finally coming into alignment. 杨斌 argues that after DeepSeek, 30B–70B models became the sweet spot for a wide range of applications, with both capability and relative cost crossing the usability threshold. As the industry moved from pilots and demos to scaled deployment, power, cost and latency stability shifted from negligible line items to sources of margin: “If your technology arrives far ahead of the industry cycle, you become an early casualty.”
  • The inference market will not simply reproduce the winner-takes-all structure of training; the long tail across edge and near-edge applications will support multiple architectures, sizes and vendors. Wearables, cameras, cars, edge appliances and embodied AI make different trade-offs among performance, power and cost, so no single “hexagonal warrior” can cover them all. Chinese companies’ structural advantages come from proximity to consumer-electronics supply chains and customers, as well as their ability to iterate rapidly across multiple dimensions.
  • The technological dividing line is that CPUs, GPUs and NPUs remain evolutionary branches of the same technology stack, while the LPU breaks with shared-memory-dominated data exchange through a deterministic hardware dataflow. The move from CPU to GPU shifted computation from scalar to matrix and vector operations; the episode characterizes the NPU as an optimization of the data-transfer paradigm. The LPU works more like an assembly line or conveyor-belt sushi, keeping data moving in a fixed direction and trading flexibility for high utilization, low power and predictable latency.
  • 源创微’s bet is “LPU Plus”: a near-term redesign of conventional hardware with embodied AI as the long-term destination. 杨斌 says the lab’s cost and latency results are “extremely close” to Groq’s data, but the product will not copy Groq; it will layer on more than 20 years of experience in large-chip design, software and markets. His startup trigger came after reading the DeepSeek paper over the 2025 Lunar New Year holiday and realizing that “large models are really not a bubble—they’re usable now,” leading him to pursue something “difficult, but right.”

Deep dive

1. The $20B Deal Is a Bet on the Industry’s Second Half

  • 杨永成 first laid out the market’s standard explanations: Nvidia may have acted to kill a potential competitor, or to acquire Groq founder Jonathan Ross for his TPU background. 杨斌 prefaced his response as a personal view, then said the first explanation was “relatively narrow.”

  • The deal size was his starting point: roughly $20B according to the episode, versus Nvidia’s approximately $56B in cash flow in the first 3 quarters of the year. On a conservative full-year estimate, the transaction could equal one-quarter to one-third of annual cash flow. “This is drawing on the family assets. It’s a lot of money.”

  • 杨斌 placed the deal in the context of the AI industry cycle: the first half competed on model parameters, context and benchmarks; the second half is about applications capturing the value of those models. Nvidia’s choice of an inference company signals that it sees inference as the next core source of compute demand.

  • 杨永成 summarized the implication: Nvidia recognizes the LPU route and expects the inference market to ramp rapidly. It wants to establish a lead comparable to its position in training and put other players “far behind.”

2. Groq Fills Nvidia’s Real-Time Inference Gap

  • 杨斌’s map of the stack is the Grace CPU, Rubin GPU and the DPU acquired from Mellanox. Together, they cover cloud-scale supernode training, but the move toward edge devices, autonomous driving and robotics still leaves a gap in real-time inference.

  • Starting in 2025, 杨斌 observed Nvidia gradually shifting from training toward inference and treating embodied AI as a second growth curve. Groq’s LPU is therefore not an isolated acquisition; it “completes the entire competitive map.”

  • For Groq, joining Nvidia offers another chance to grow and potentially strengthen its capabilities. Advanced process technology, packaging, IP, product integration and engineering resources can accelerate iteration on the architecture. The signal visible later in the episode was that the core team joined Nvidia’s product development efforts, with the LPU positioned as a foundation for “real-time inference.”

3. Revenue Shows Groq Is More Than an Architecture Experiment

  • 杨斌 gave the operating figures as roughly $90M in revenue for Groq in 2024 and an expected approximately $500M in 2025. He added that he did not know whether the current figures had changed, then characterized Nvidia’s move as “a very mature decision, not a very impulsive one.”

  • Groq’s customers fall into 3 groups: US tech giants such as Meta; national compute centers in Saudi Arabia and Norway; and retail users and small businesses on the cloud. Its cloud service has around 2.5M registered users.

  • That also weakens the “hostile acquisition” interpretation. Being treated as a competitor by Nvidia is itself a strong validation of Groq, but the episode puts greater weight on the combination of revenue growth, real-time inference capability and Nvidia’s strategic needs.

4. The “TPU Talent Acquisition” Theory Fails the Timeline Test

  • 杨斌 acknowledged that Jonathan Ross was the chief architect of TPU V1 and is often called the “father of TPU.” But that was in 2016; he subsequently left Google and independently developed Groq’s LPU, nearly a decade ago.

  • His counter-question was straightforward: if the goal were merely to acquire TPU talent, Nvidia could go directly to Google. 杨永成 added that TPU has since evolved to V7 and changed substantially from V1. More importantly, TPU serves both training and inference, while LPU is an entirely different “species.”

  • 杨斌 offered another possible explanation, while stressing that it was only his interpretation: separating the IP, team and brand may relate to US antitrust constraints. 杨永成 compared it to Intel needing AMD to remain in existence—the Groq brand and cloud business can continue operating independently while key personnel join Nvidia.

5. After 9 Years, the LPU Found Its Moment in a Technology-Industry Convergence

  • 杨斌’s warning was blunt: “When your technology arrives far ahead of the industry cycle, you become an early casualty.” Groq was not a research institution; however strong the technology, it still needed a commercial loop. Inference volumes were previously too small, and GPUs were already good enough. Building an entirely separate system simply did not pencil out.

  • The inflection point came when model capability and cost crossed the threshold together. The episode argues that after DeepSeek drew broad attention, 30B–70B models became the sweet spot for a wide range of applications: “Basically, we can do everything we can imagine.”

  • 杨永成 described the earlier phase as pilot-style applications focused on proving that something could be done and enjoying the novelty. Waiting another fraction of a second or accepting somewhat higher power consumption was tolerable. Once applications entered millions of households, cost, electricity, latency and user experience all had to be optimized, allowing an inference-specific LPU to show its value.

  • The two men summed up the timing with a proverb: “The early bird gets the worm,” but it can also be “the early worm gets eaten by the bird.” Had the inference market arrived 5 years later, Groq might still have been an insightful technology that was difficult to monetize.

6. Stable Low Latency Determines Whether a Use Case Works

  • 杨斌 stressed that the LPU’s appeal is not only low absolute latency, but its ability to sustain low latency over time. When people talk to robots, a slow response breaks immersion; an erratic response disrupts the rhythm of the interaction just as much.

  • 杨永成 extended the logic to autonomous driving: if compute output is fast one moment and slow the next, vehicle control becomes unacceptably risky. Deterministic latency is therefore fundamentally a question of whether the system works at all, not merely whether the experience feels good.

  • The consumer-electronics example was blinking in photos. Today, a PC or high-end phone can post-process a closed-eye shot and turn it into an open-eyed image. If real-time inference is fast enough, the camera can generate an open-eyed, smiling version as the shutter fires, converting post-processing capability into a feature that creates value directly.

7. The 14nm-versus-4nm Comparison Shows the Architecture Dividend

  • 杨斌 cited data Groq presented at Hot Chips 2024: compared with Nvidia’s 4nm H100, Groq’s own 14nm architecture delivered roughly one-sixth the latency, one-third the power consumption and one-quarter the cost, with Groq claiming 10x higher overall energy efficiency.

  • Those metrics are particularly relevant for edge and consumer products, where cost, power consumption and real-time experience all matter at once. 杨永成 joked that this was already enough to build his “dream camera.”

  • 源创微’s laboratory results on cost and latency were “extremely close” to Groq’s. 杨斌 said the team was not stopping at replication but had carried out extensive new iterations: “This order of magnitude can definitely be monetized,” and he is confident the final product can do somewhat better.

8. The LPU Is a “New Species” Grown Outside the Shared-Memory System

  • 杨斌 started with the von Neumann architecture: the CPU uses shared memory as the exchange center, trading generality for instruction-driven, programmable data flows. It can run AI workloads, but not efficiently.

  • The GPU does not change the underlying data-exchange model; it extends the CPU’s scalar computation into matrix and vector operations, producing Tensor Cores and Vector Cores. 杨斌 then used the addition of “special channels” between processors to describe one class of transfer optimization, noting that NPUs are broadly similar to GPUs without going further.

  • The LPU removes public shared memory and sends data along a fixed direction, creating a deterministic hardware dataflow architecture. 杨斌’s analogy was a banquet where the dishes never move versus conveyor-belt sushi: “What you want is right in front of you.”

  • 杨永成 offered an organizational analogy: a GPU is like a professor leading a group of PhDs who can all do different jobs—flexible, but with scheduling overhead. An LPU is like an assembly line, where each station performs one task and the product moves continuously from one end to the other, making efficiency and latency more controllable.

9. Extreme Inference Performance Requires Giving Up Training Flexibility

  • 杨斌 did not say the LPU cannot train models. Groq initially “forced itself to do training” and still operates cloud business in Saudi Arabia. But training paths and model architectures change quickly, and the LPU is less flexible than the GPU.

  • Once a model stabilizes, the LPU’s efficiency and real-time advantages re-emerge; when models change frequently, they may need to be recompiled. The trade-off is explicit: 源创微 has “abandoned training” to push inference to the limit.

  • That specialization also means no single chip will cover every market. Other vendors may sacrifice performance to optimize power consumption and still find customers. “There cannot be one unique hexagonal warrior solving every problem.”

10. The Real Barrier Is the Complete System, Not Just a Chip Diagram

  • Multicore CPUs and GPUs can replicate large numbers of identical cores. Every stage in an LPU pipeline is different, however, and must be designed, verified, optimized and connected separately. The hardware workload is therefore not a matter of simply adding more cores.

  • The software challenge is equally significant. Customers want to keep using familiar model libraries and development workflows, while the compiler must understand an entirely new non-von-Neumann architecture. A complete team that understands processors, compilers and AI architectures is “extremely rare.”

  • Groq, as a pioneer, has already worked through many of the pitfalls, but its early lack of market feedback constrained iteration speed. The episode also mentioned industry criticism of Groq’s heavy use of SRAM, involving power consumption, cost and die area.

  • 杨永成 believes the opportunity for Chinese teams lies not in copying a giant’s design, but in redesigning around proximity to consumer-electronics customers and accumulated technical, mass-production and demand-side feedback. That is why 源创微 calls its approach “LPU Plus.”

11. China’s Opportunity Lies in Long-Tail Markets and Conventional-Hardware Redesign

  • 杨斌 sees inference as a relatively blue-ocean market that will turn red quickly once demand scales. The edge spans wearables, cameras, cars, edge appliances and embodied AI, while different model sizes and power constraints will create room for multiple architectures and companies.

  • China’s structural advantages are closer access to supply chains and customers, along with greater strength in high-frequency, multidimensional iteration. The long-term “stars and sea” is embodied AI; the near-term entry point is “the redesign of conventional hardware”—turning classifiers that only recognize cats, dogs or falls into systems capable of generating content.

  • 杨斌’s decisive moment as an entrepreneur came during the 2025 Lunar New Year holiday. After reading the DeepSeek paper, his first reaction was: “Large models are really not a bubble—they’re usable now.” After surveying paths including GPUs, NPUs and LPUs, he finally concluded that the application-boom cycle had arrived.

  • The more-than-RMB100M angel round also attracted industrial capital from 2 listed companies. 杨斌 summarized the final choice as an LPU Plus that combines the strengths of multiple approaches, but he places greater weight on one conclusion: “The LPU is difficult, but it is the right thing to do.”