28. Inside GTC 2025 (Part I) | Nvidia's History: How One Company Kept Betting on the Future for 30 Years
Summary
- Nvidia’s dominance was not the result of a single winning AI bet, but of turning future applications into platforms over 20-30 years—from graphics GPUs and GPGPU to CUDA and supercomputing clusters. 姚欣 believes the hardest thing to replicate is Jensen Huang’s “patience with the vision”: a decade ago, he was still personally promoting Tegra from company to company; in 2024, about 80% of Nvidia’s workforce was working on software, while demos were seeding future demand for compute in the metaverse, quantum computing, humanoid robots and other areas.
- Nvidia remains in a league of its own in training, but inference will likely evolve from a monopoly into a market with one dominant player and many specialists—the central structural turning point for its valuation. 姚欣’s framework: Nvidia could hold more than 50%-60% of training, while a 20% inference share would already make it a strong leader; as specialized chips develop, it could also fall back to being an ordinary participant with 5%-10% share, because the “impossible triangle” of cost, performance and power consumption will generate a large number of specialized solutions.
- Nvidia’s near- and medium-term revenue and its long-term share price can move in opposite directions: revenue could still meet or exceed investor expectations over the next 2-3 years, but at least one of its high gross margins and monopoly position will decline. 姚欣 describes the current setup as “profits are good, growth is good, and the market is monopolized,” and says that this three-way combination “will never last long in any period of technology history”; AMD, in-house XPUs and Chinese supply are unlikely to create a more meaningful challenge until 3-5 years from now.
- DeepSeek did not make compute unnecessary; it shifted incremental demand from training toward inference, forcing the market to recalculate the cost, memory requirements and delivery cadence of each inference run. 姚欣 is watching to see whether the new card reaches 288GB of memory, whether it also has IP4 capabilities (not elaborated in the original), and whether one machine can run the “full-strength version of DeepSeek”; if deliveries begin by year-end, Nvidia’s inference share could return to 60%-70%, but a delay until June next year would be “below expectations.”
- The campaign against Nvidia is not being waged by one rival, but by giants that are simultaneously buying GPUs, backing AMD and ROCm, developing specialized chips such as TPU, and investing in startups. AMD, the main traditional competitor, “has all the capabilities, but not the ecosystem”; Google has built a small closed loop around TensorFlow, TPU, the cloud and internal demand, Broadcom is taking the custom-chip business, and 姚欣 believes only Huawei can currently discuss competition with Nvidia in selected domestic markets.
- Jensen Huang is actively reducing dependence on major customers by combining chips, capital and financing leverage into a second demand network. From Oracle Cloud to CoreWeave as a “white glove” intermediary and then to “sovereign AI” for governments around the world, the logic is to prevent a handful of hyperscalers from absorbing most of the demand; when 卫诗婕 asked whether this had become a financial-capital game, 姚欣 compared it directly with “the next-generation LeEco,” suggesting that the strategy combines customer-mix optimization with short-term excess.
- Higher inference efficiency will not automatically end compute growth, because cheaper inference will expand usage; images may consume roughly 1000x as much compute as text, and video could add another 1000x. 姚欣 estimates that inference could rise from about 40% of compute in 2024 to more than 95%, while multimodality could expand the total market by “1000x or 10,000x”; Nvidia will still help “build the infrastructure for the next generation of human civilization,” but “it will not be the only one.”
Deep dive
1. Nvidia kept rewriting what GPUs were for, rather than winning AI with one bet
卫诗婕 opens with the idea that “the person selling the shovels is happiest,” then asks how Nvidia evolved from a graphics-card company into the infrastructure provider of the AI era.
姚欣 starts with the era of SGI, Lucas and Star Wars: graphics computing first served animation and film special effects, after which Nvidia entered the market with graphics cards and then pushed graphics GPUs into GPGPU, or General Purpose GPU.
The clearest evidence of the platform strategy is the steady migration of applications: more than a decade ago, PPTV used CUDA for audio and video codecs; then came cryptocurrency mining, deep learning in 2013, AlphaGo and Omniverse. “One towering tree after another grows atop a general-purpose platform.”
2. Jensen Huang builds tomorrow’s demand into today’s demos
姚欣 believes Nvidia’s real hidden capability is Jensen Huang’s vision and his “patience with the vision”: many of his calls take 10 or 20 years to mature, requiring him to remain patient and lie in wait for the long term rather than validate everything within a single product cycle.
Tegra is both the counterexample and the best illustration of that patience. The chip performed well but consumed too much power; Huang personally promoted it for vendors including Xiaomi tablets, but the mobile strategy gradually lost to ARM and Qualcomm. A decade later, GPUs determined which companies could train the largest models, producing a new Silicon Valley divide between “GPU rich” and “GPU poor.”
In early 2024, Nvidia told 姚欣 that about 80% of its people were working on software—not just CUDA, but also applications built on top and a large number of demos. The aim was to show the entire industry why the metaverse, quantum computing and humanoid robots of 10 years from now already need Nvidia GPUs today.
3. CUDA and clustering turned chip advantages into system-level moats
The first layer of Nvidia’s moat is mature, best-in-class hardware; the second is CUDA. 姚欣 likens CUDA to “the C or Basic of the AI era”: after more than 10 years of training, developers in AI, graphics rendering and image processing have built up deep habits, making migration costs far higher than simply replacing a card.
The 2019 acquisition of Mellanox was the key inflection point in training. Nvidia moved from selling individual GPUs to selling supercomputing clusters made up of thousands of cards, using TB-scale networking to solve communication across the cluster. Both the commercial unit and the technical moat expanded at the same time.
The failed ARM acquisition also defined Nvidia’s limits. Masayoshi Son pushed for a combination of ARM and Nvidia that would cover the entire chip stack from phones to the cloud, but the deal was blocked by governments in multiple countries. Tegra subsequently faded, and Nvidia largely abandoned the mobile battlefield.
4. Silicon Valley’s “godfather” was made by the era, academic networks and organizational design
姚欣 rejects excessive deification of Jensen Huang: “Five years ago, he was a nobody in Silicon Valley.” Eight years ago, few people wanted an Nvidia offer; around 2022, entrepreneurs saw selling to Nvidia as a full tier below selling to Meta or Google. AI demand turned him into a hero of the era.
姚欣 sees the leather jacket, Stanford’s Huang Building and GTC’s long-term cultivation as expressions of Huang’s strong instinct for branding. More than 10 years ago, GTC was an academic forum and workshop for professors, students and developers; today it opens with rock music and resembles a major concert, while Nvidia has continued cultivating the same technical communities for more than a decade.
The academic network matters more than the personal mythology. 姚欣 cites Stanford professor Bill Dally, who later became Nvidia’s chief scientist, and argues that the GPGPU standard rests on the foresight, engineering ability and influence of university professors. Huang’s strength is integrating those resources into the company.
Around 50 people report directly to Jensen Huang. 卫诗婕 believes the no-PowerPoint “whiteboard culture” allows people doing the frontline work to be seen and prevents seniority from dominating; 姚欣 adds that Huang cannot cover everything himself, but can stay informed about the frontline, emerging ideas and innovative startups.
姚欣 also cites a Silicon Valley AI search company whose name sounded like “Papalaxy” in the original: Huang personally tested its product, met its 10- to 20-person team, and supported and invested in it. The episode shows that even under extreme time pressure, he has remained sensitive to new ideas.
5. Taiwan’s hardware network is Nvidia’s hidden supply-chain foundation
Nvidia depends not only on TSMC’s advanced capacity but also on competing with customers such as Apple for supply. 姚欣 believes Huang has helped integrate Taiwan’s semiconductor and server-hardware network by following Taiwan’s technology and electronics trade shows and bringing along partners such as Super Micro.
Super Micro integrates GPUs, memory and storage onto server motherboards. 姚欣 says the business “doesn’t have particularly deep technical value,” but Nvidia’s compute demand once drove its stock up roughly 10x, showing how growth can spill over across the supply chain.
From the 2000s through roughly 2012-15, more than 90% of the assemblers and manufacturers of core computer components were effectively based in Taiwan; factories might have been located in mainland China, but the brand operators were mostly Taiwanese. Combined with the early Silicon Valley network of Chinese semiconductor professionals, this “Taiwanese hometown network” formed an underappreciated global manufacturing base for Nvidia.
6. Intel’s decline proves that if you do not compete for the future, the future will crush you
After 2011-13, CPU Moore’s Law began to slow just as new applications were demanding more parallel computing. That was the technological backdrop for the GPU’s structural opportunity.
Intel was also trapped in the “innovator’s dilemma”: the high profitability of its PC business kept it from betting sufficiently on mobile, eventually ceding the market to ARM and Qualcomm. Nvidia missed mobile as well, but retained another growth curve in general-purpose GPUs.
The manufacturing slowdown was more damaging. 姚欣 says Intel remained at 14nm for roughly 6 years, merely “changing the name and the shell,” while TSMC and Samsung advanced to 7nm and 5nm. After AMD abandoned its own manufacturing and embraced advanced foundries, it overtook Intel in process technology and energy efficiency. “If you don’t compete for the future, one day the future will compete you to death.”
7. Training remains in a league of its own; inference is destined for one strong player and many specialists
姚欣 draws the market-share lines this way: more than 40% makes a monopolist, around 20% makes a leader, and 5%-10% makes a participant. By that standard, Nvidia may hold more than 50%-60% of the training market, with virtually no rival capable of challenging it head-on.
Inference follows specific applications and must make trade-offs within the “impossible triangle” of cost, performance and power consumption. The cloud prioritizes performance, while glasses, cameras and automotive chips care more about energy efficiency; as in the mobile era, when the market chose ARM, different use cases will choose different architectures.
The more likely end state for inference is one dominant player—or one or two dominant players—alongside many vertical specialists. Nvidia may still hold 40%-50% in the near term; over time, even a roughly 20% share would make it a successful leader, while the rise of specialized XPUs could reduce it to one participant among many.
8. The giants are using every form of competition against Nvidia
AMD is Nvidia’s main traditional competitor and the only company with both x86 CPU and GPU capabilities. The problem is that fighting on multiple fronts disperses its investment, while its software ecosystem is far less developed than CUDA’s, leaving it in the position of “everyone’s consensus favorite that no one can carry.”
ROCm is trying to become the “Android of the AI-chip era.” It is open source, backed by supporters including Meta and Microsoft, and intended to reduce dependence on a single hardware platform; Nvidia and CUDA occupy the corresponding “iOS” position.
姚欣 describes the giants’ strategy as “pretty fragmented”: on one hand they “flatter Nvidia” and continue buying GPUs; on another they commit heavily to AMD; on a third they develop their own chips, while also incubating and investing in small teams. “They are using almost every form of competition available.”
Broadcom is another beneficiary of the challenge from specialized chips. It builds proprietary chips customized for specific industries; 姚欣 sees its entry into the trillion-dollar market-cap club as a signal that investors are betting Nvidia will not remain the industry’s only dominant player.
9. Google uses a closed internal loop to keep TPU’s seat at the table
In the AI 1.0 era, TensorFlow sat on top of CUDA while also calling Google’s in-house TPU. Developers did not need to understand the underlying layer; they only needed to use Google AI Cloud. Google’s own enormous usage then fed back into TPU iteration, capacity optimization and continued investment, creating a small vertical loop.
姚欣 believes AWS and Microsoft will have a harder time replicating the model: enterprise customers demand transparency and clarity, and will specify Nvidia GPUs; Google’s internal businesses and smaller customers are more willing to accept a black-box solution. But with the second AI wave driven by Transformer, the “temporary entry ticket” Google obtained around 2016 is no longer as powerful as it was.
10. Chinese chips are losing on capacity for now; the landscape may change only in 5 years
姚欣’s blunt assessment is that only Huawei’s Ascend can currently compete with Nvidia in selected domestic niches. Chip projects at other internet companies have yet to build meaningful scale; the best chances belong to very large cloud operators such as Alibaba, Tencent and Huawei.
The bottleneck is not just chip design. 姚欣 estimates process technology may account for only 20%-30% of the problem, with capacity accounting for 70%-80%. He notes that some people in the market believe 5nm and 7nm are already sufficient for today’s AI, especially inference, but stresses that high-end capacity is not even sufficient for domestic needs, let alone exports.
Chinese chips at 28nm and above reportedly already account for about 45% of global sales, while 14nm could eventually represent half of the world market. If capacity shifts from shortage to abundance as it did in new energy, Chinese chips could “sweep across the field” in multiple industries 5 years from now.
11. Crisis, revenue growth and a valuation selloff can all happen at once
On the claim that Nvidia was “30 days from bankruptcy,” 姚欣 says it is exaggerated, although Nvidia has indeed faced danger several times in its history. Because the founder owns an unusually small stake, the company is fundamentally an investor-controlled public company; Huang’s longevity is tied to correctly catching a new wave every 3 or 5 years and laying out ahead of it.
姚欣 believes 80%-90% gross margins cannot persist: “There has never been any kind of infrastructure in any era” that could continue developing at such margins. Excessive monopoly may already be constraining industry progress, turning Nvidia into a target for everyone.
He is highly pessimistic on the long-term stock price, calling the idea that high gross margins can support 30 years of future growth “a colossal lie.” He compares Nvidia’s three advantages—profits, growth and monopoly—with the Davis double whammy/double boost, arguing that either margins or monopoly power will gradually decline.
卫诗婕 summarizes the moment as “the weight of the crown”: Nvidia enjoys its position as king while absorbing attacks from every direction. 姚欣 agrees, explaining why Jensen Huang must keep talking about humanoid robots and other third growth curves—the story of 3-5 years from now cannot be proved or disproved today.
The time horizons cannot be conflated. Over the next 2-3 years, challengers still have to climb the three mountains of capacity, performance and ecosystem, and Nvidia’s revenue can still meet or exceed investor expectations; margins will fall, however, and its market share and architectural advantage will face a more serious challenge in 3-5 years.
12. DeepSeek moved GTC’s decisive variables from compute to memory and delivery
DeepSeek’s short-term shock is that the shift from training to inference left inference usage far below the previous linear expectations, creating a demand gap. 姚欣 also believes falling prices will broaden inference adoption and that scale growth could eventually fill the hole.
Inference typically does not require a massive cluster; 1 or 2 servers may be enough. But if the full model cannot fit into memory, it must be scheduled across cards and machines. I/O throughput and scheduling across memory, rather than the compute chips themselves, are the key reasons inference performance is failing to scale today. That is why 姚欣 is turning his GTC focus to memory capacity.
姚欣 is watching whether the new card reaches 288GB and whether it also has IP4 capabilities (not elaborated in the original). His team’s calculations suggest that if both conditions are met, one machine could run the “full-strength version of DeepSeek,” driving inference costs lower and potentially restoring Nvidia’s initial share to 60%-70%.
But specifications are not earnings; shipment timing determines the market reaction. GB200 was launched in March last year and still had not fully arrived by the time of the conversation, while the first deployments had defects; Nvidia even reduced the GB200 allocation within its TSMC capacity. His conclusion: “There is no harm in talking big, but after you talk big, you have to deliver.” Otherwise, the next vision will first be marked down to 80%, then to 60%.
13. Nvidia is rewriting its customer mix with “white gloves” and sovereign AI
Nvidia was still “the good kid selling chips” in 2017, but took a major hit in 2018 as the mining boom faded and alternatives such as Google’s TPU and FPGAs emerged. The company realized that selling 80%-90% of its output to a handful of giants was tantamount to “playing with a tiger.”
Starting in 2023, Nvidia restricted supply to the largest players while supporting second-tier clouds. Oracle Cloud may have held less than 5% of the US market, and perhaps less than 3%, yet received supply at roughly the same scale as Google, Microsoft and Amazon. Nvidia used the arrangement to “build out Oracle’s skeleton.”
The CoreWeave chain is more aggressive: Nvidia invested in CoreWeave; CoreWeave used the funds to buy Nvidia GPUs, then pledged the cards to banks for roughly 80% financing to keep buying more. 姚欣 says the company grew from a small New Jersey mining operation into a business valued at about $25B in 3 years, while Microsoft and OpenAI were then forced to rent its cards.
When 卫诗婕 asked whether this had become a financial-capital game, 姚欣 answered, “It’s the next-generation LeEco.” At the same time, Nvidia is selling domestic-language models and “sovereign AI” to governments in the Middle East, Europe and elsewhere, creating new, state-funded demand outside the handful of major clouds.
14. Efficiency gains will create only a temporary hole; multimodality will expand the compute market
DeepSeek led the market to make a seemingly reasonable estimate: if thousands of servers and tens of thousands of cards can serve 100M+ users, multiplying that across 50-60 to 100 large apps worldwide suggests Nvidia’s current annual supply of roughly 1.2M cards—and a potential future supply of 3M-5M—could already be enough.
姚欣 believes that estimate misses the upgrade in modalities. Image generation may consume 1000x as much compute as text, and video could add another 1000x; looking back to the early internet, he notes that video storage and transmission could be 100x larger than audio, while the server consumption of audio and the text-and-image internet occupied yet another order of magnitude.
Inference could therefore rise from about 40% of compute in 2024, through 60% and 70%, to more than 95%, while the total market could expand by 1000x or 10,000x. The near-term shift in growth momentum can hit the stock price, and Nvidia may still have no mature rival in the medium term; over the long term, compute demand will continue to explode, but Nvidia will be only one of the builders of that infrastructure—not the only one.