E230 | Behind the $1T Revenue Expectation: Nvidia's Peak and Soft Underbelly
Summary
黄仁勋’s core expectation is that cumulative orders for Blackwell and Vera Rubin will reach at least $1T by the end of 2027, double the previous $500B benchmark. That is even higher than 2024 global semiconductor sales of more than $600B, but 张璐 believes the demand logic holds: training is a one-off investment, while Agent inference generates recurring cash flow; the cost mix could shift from training at 70%-80% in 2023 to inference at 70%-80% in the future—“Long-term cash flow must come from inference” (“长期的现金流,那一定是来自于推理”).
The hardest hurdle to the $1T expectation is not orders but physical capacity that capital cannot buy on demand. 肖志斌 sees 3nm wafer capacity as likely to keep up, while CoWoS advanced packaging is harder to call; capacity has grown roughly 3x since 2024, Micron and Samsung announced HBM2 mass production in March, and Micron, Samsung and SK hynix are advancing customized HBM4E solutions, but production-line investment, process optimization and validation still take 1-2 years. “This isn’t software: however much demand you have, you can’t instantly turn it into an equivalent amount of sales.”
NVIDIA’s moat has expanded from CUDA into full-stack execution spanning chips, systems, supply chain, developers and customer feedback. Vera Rubin launched 7 chips simultaneously, all already in mass production; NVL72 delivers 10x Blackwell’s inference efficiency, cuts token cost to one-tenth and raises token per watt 35x. Mark says NVIDIA internally went from no one using coding agents to “100% using them” within 1-2 months, but the real challenge remains turning generated designs into end-to-end-optimal designs.
Groq’s LPU shows that inference will not be dominated by a single chip type; future data centers are more likely to be heterogeneous systems combining GPUs, LPUs, Switch and optical interconnect. Groq uses on-chip SRAM to store weights and KV Cache, sacrificing capacity and cost for ultralow latency, especially for token-by-token decoders and Agents; 黄仁勋 even suggested reserving 25% of data-center space for Groq and other inference chips. 肖志斌’s blunt view: “GPUs are actually not very well suited to agentic applications.”
TPU is a real threat, but not yet enough to overturn NVIDIA’s position as a third-party infrastructure provider. Google can vertically optimize models, chips, interconnect, power and applications, but external customers may not replicate its internal efficiency; NVIDIA relies on cross-customer software, exceptional execution and control of TSMC and CoWoS capacity to defend its lead. The more consequential gaps are private deployment, Edge AI, robotics and specialized CPUs/NPUs, while its huge market cap means it is “celebrated by capital, but also held hostage by capital.”
OpenClaw and NemoClaw are pushing token demand toward always-on Agents and shifting SaaS’s unit of competition from software seats to AI labor. 张璐 believes vendors may move from IT budgets into the larger labor budget, but using hiring standards as a benchmark, she argues an Agent would need to complete more than 90% of a role and outperform more than 90% of people to truly replace them; there is still a gap. SaaS companies without model capabilities that fail to transform quickly face the highest risk. “Future software companies … become labor suppliers.”
Data centers have already validated extremely strong demand, but $1T in revenue ultimately depends on land, power, delivery speed and operational reliability. Alex says roughly 90% of new US projects are shifting to behind-the-meter natural-gas generation, while modular construction cuts the time from greenfield to live service from 18-20 months to 6-9 months. Memory prices are up 100%-200%, shortages are emerging in SSDs, ConnectX-7, switch gear, Intel CPUs and CDU liquid cooling, and CX7 is also shifting to BlueField; the supply chain is unlikely to ease meaningfully by the end of 2027. “Capacity exists, but the price is impossible to determine.”
Deep dive
1. The $1T Is a Revenue Forecast for an Entire AI Factory, Not a Single-GPU Story
泓君 framed this GTC around 4 figures: cumulative Blackwell and Vera Rubin orders of at least $1T by the end of 2027; 7 chips launched simultaneously on the Vera Rubin platform; NVL72 delivering 10x Blackwell’s inference efficiency and cutting per-token cost to one-tenth; and token per watt improving 35x.
That figure doubles last year’s $500B benchmark and resets the comparison. 2024 global semiconductor sales were only more than $600B, while Lisa Su had previously discussed data-center AI accelerator chips reaching $1T by 2030; 黄仁勋 pulls the timeline forward to 2027 and concentrates the scope in NVIDIA’s complete system alone.
肖志斌 points out that Vera Rubin sells far more than chips: it also includes NVLink Switch, Ethernet Switch and software. 张璐 therefore redefines NVIDIA as an “AI infrastructure company,” whose AI factories produce not compute itself but the unit in which future productivity is measured—the token.
The explanation from public-market investors at the event is worth preserving: the stock showed no obvious reaction after the $1T figure was announced, not because the number was small, but because “many Wall Street investors had already modeled roughly this number when they built their models.”
2. Inference Moves from Cost-Side Supporting Role to Recurring Cash Flow, the Core Logic Behind $1T Demand
张璐 distinguishes training from inference: training is closer to a one-off capital investment, while inference recurs with every Agent call. Long context, low latency, real-time availability and continuous interaction all increase token consumption, so “if you look at long-term cash flow, it has to come from inference.”
She recalls that training may have accounted for 70%-80% of model costs in 2023; today the split may be close to 50/50. Over the next 1-2 years, inference could instead rise to 70%-80%. This is not a firm forecast: the pace depends on how quickly industries integrate AI and deploy Agents.
黄仁勋 offers another set of figures: inference compute has increased 10,000x over the past 2 years and usage has risen 100x, implying a 1M-fold increase in compute demand. Alex sees the corresponding signal in GPU cloud: “one cluster, basically, has 3-4 hyperscalers competing for it.”
3. 3nm May Keep Up; CoWoS, HBM and Process Cycles Are the Hard Supply Constraints
肖志斌 sees the $1T order figure as a “very strong” demand signal, but says delivery by 2027 will be challenging: 3nm wafer capacity will probably keep pace, while CoWoS advanced packaging is harder to call, even though TSMC has expanded CoWoS capacity by roughly 3x since 2024 and continues to add more.
Memory makers are also playing catch-up. 肖志斌 says Micron and Samsung announced HBM2 mass production in March, while Micron, Samsung and SK hynix are working on customized HBM4E solutions. The problem is that demand growth pulls simultaneously on wafers, packaging, memory, switches, testing and power; there is no shortcut in expanding just one link.
张璐 emphasizes the fundamental difference between semiconductors and software: production lines require upfront investment, process control and yield optimization, and incremental demand often takes 1-2 years to become capacity. “You can’t buy your way through this cycle”; by then, product requirements may already have changed.
4. 7 Chips in a Year Requires More Than Hiring: AI and Closed-Loop Feedback Compress the Design Cycle
Mark says NVIDIA moved from roughly 1 chip every 2 years to 1 per year, and then to multiple chips per year. The traditional answer was to expand the team, but coding agents and AI for chip design have materially lifted engineer productivity; the number of chips is not the hardest part—AI is needed to improve the quality of on-chip optimization.
Traditional chipmakers hand products to customers and wait for feedback before iterating, creating long communication cycles. NVIDIA has CUDA, enterprise customers and a startup ecosystem at the same time, allowing it to observe system bottlenecks directly and quickly select the top 3, top 5 and top 7 priorities from 20 possible actions.
That closed loop, combined with supply-chain coordination, lets chip design, demand assessment and mass-production preparation proceed in parallel. Against the traditional semiconductor cadence, shipping 1-2 chips a year is already excellent; Vera Rubin’s 7-chip launch is about system-level execution, not just point-design speed.
5. Groq Trades SRAM for Low Latency and Hits the Communication Bottleneck in Agent Inference
肖志斌 explains that SRAM access latency is roughly 1-2 nanoseconds and requires no dynamic refresh, but each storage cell needs about 6 transistors. DRAM uses roughly 1 transistor per cell, giving it higher density and lower cost, but it comes with higher access latency and a refresh burden.
Groq takes an unconventional route: it removes external DRAM, stores model weights and KV Cache in on-chip SRAM, and relies on extreme interconnect to scale clusters. That eliminates the need to fetch weights from memory again for every token and makes token per second per user more stable; based on the curve 黄仁勋 showed, efficiency in specific Agent applications can be more than 30x that of GPUs.
In inference, the encoder fits the GPU’s strength in high-throughput batch processing, while the decoder generates one token at a time, with little compute but heavy communication. 肖志斌 puts it directly: “Most of the time is actually spent fetching weights—communication, not compute.”
张璐 folds energy use into the same logic: compute power consumption continues to fall, while communication power declines more slowly. She relays John Hennessy’s view that future communication power draw could reach more than 10x compute power; the LPU’s value is not only speed, but also reducing data movement.
6. Inference-Chip Startups Still Have Openings, but Directly Replicating NVIDIA Is a Poor Bet
张璐 does not say the startup window is closed, but believes the opportunity for a general-purpose inference chip is relatively small. A more realistic path is to find gaps outside NVIDIA’s top 3-to-7 priorities and become a partner, an ecosystem patch or a potential acquisition target rather than collide with NVIDIA’s resource deployment head-on.
Her example is the interconnect switch. The team originally studied optical compute, but after learning more about NVIDIA’s internal progress, it shifted toward next-generation interconnect. The bottleneck in large-scale AI infrastructure is spilling from compute into switches and interconnect, an area that is both competitively adjacent to NVIDIA and strategically complementary.
肖志斌 uses 2 startups to revise his own judgment. In 2017, he built a pure-SRAM inference chip at Alibaba; it ranked first in the world on MLPerf, but when applications changed, the team could not keep up with the GPU software ecosystem. Later, DeepSparse combined model compression with chip architecture and once compared 4 chips against 1 H100 on BERT, but the large-model era left it bottlenecked by expensive post-training.
He has now shifted to neutral, system-level optimization: emulating Google TPU, AMD GPU and NVIDIA systems, leaving kernels to chipmakers and doing automated optimization at the upper layers. Mark compares it to the cloud-computing revolution: every shift in computing paradigm opens many narrow but deep startup positions across the stack.
7. OpenClaw Brings a Token Surge; NemoClaw Is Really Contesting Rule-Setting Power over Enterprise Agents
肖志斌 describes OpenClaw’s potential incremental demand as the “1,000x token usage” 黄仁勋 discussed. He mixes models to control costs, relying primarily on Kimi 2.5, which is about 10x cheaper than Claude; his company plans to open-source Token Simulator and Auto Optimize, which will automatically plan model calls based on a user profile.
肖志斌 believes OpenClaw benefits vertical Agents. Large companies have built substantial GPU capacity and need to use the cloud to capture usage and turn tokens into applications; an open ecosystem will also force legacy tools to become agent-friendly, allowing vertical teams such as chip designers to avoid rebuilding general-purpose tools.
张璐 does not see NVIDIA launching NemoClaw mainly to capture application revenue. It is more about establishing rule-setting authority over Agent deployment, security access and enterprise rollout. She also preserves the difference in user experience: OpenClaw tends to “get the thing done,” while enterprises require it to “do the thing well.”
She personally prefers Claude Code for high-quality enterprise work and mentions the competition from Claude Dispatch. 泓君 gives the example of OpenClaw rapidly writing a CRM highly tailored to a single company’s business, exactly targeting the demand that traditional standardized SaaS struggles most to cover.
8. Agent as a Service Will Rewrite SaaS, but Replacement Speed Depends on Quality, Not Demos
张璐 describes the evolution as “foundational technology innovation—technology application innovation—business-model innovation.” Traditional SaaS sells standardized software shared by all customers; Agents can be highly customized to a company, role and workflow. What is ultimately sold may be neither software nor services, but AI labor.
The budget pool will change as well. Software has traditionally competed for IT budgets; if the output is AI labor, it could move into the larger labor budget. But using hiring standards as a benchmark, she posits that customers will require an Agent to complete more than 90% of a role and outperform more than 90% of people. Her judgment is that Agents still have some distance to go.
She disagrees that SaaS will disappear wholesale overnight: enterprise sales also include after-sales support, relationship networks and implementation capabilities. But SaaS without model capabilities will likely eventually disappear or be replaced; companies with industry data and model capabilities can still transform, while founders can capture markets left behind by incumbents’ exits.
肖志斌 reduces the transformation to 2 tasks: combine industry expertise with an Agent platform as quickly as possible, and buy and optimize compute so the ROI from compute investment into service output is maximized. “If it doesn’t make drastic changes now, it will soon be replaced by these Agent Platforms.”
9. Companies Will Manage People and Agents Together, and Token Quotas Will Become a New Organizational Resource
张璐 envisions high-value companies with only 20-30 core employees, while HR, CFO, finance and other functions are provided by project-based Agents. These Agents need not be full-time employees of a single company; organizational boundaries will shift from fixed roles to hybrid labor called in by task.
A CEO’s new capabilities will include distinguishing which departments must be human-led, which can be outsourced to Agents, and how to evaluate AI employees. 泓君 relays 黄仁勋’s hiring vision: in addition to annual salary, engineers should know how many tokens they have available and how many “Agent interns or employees” they can manage.
This also explains the rapid adoption of coding agents inside NVIDIA. Mark confirms that, starting at the beginning of the previous year, the company went from almost no usage to “100%” within 1-2 months; other chip companies have also begun similar broad deployments.
10. ChipNeMo Can Read Documents and Write RTL, but End-to-End Optimization Remains the Hard Part
Mark’s design-automation research team had been testing AI for chip design since the CNN and GNN era, but older models could solve only local problems. Language models and Agents brought more general design capabilities for the first time, helping interpret design requirements and generate code.
NVIDIA released ChipNeMo in 2023, training and adapting a base model on more than 20B internal tokens for interaction with chip-design knowledge. The first use case is a chatbot that understands internal documents and design requirements; the second is a coding agent that generates software code and even RTL hardware code.
Mark does not avoid the quality gap: “The quality is not that high yet; there is still room for improvement.” Generated RTL is only the starting point; the real difficulty is building an end-to-end design loop and optimizing the design.
11. The TPU Threat Is Real, but Google’s Internal Efficiency Does Not Equal External-Customer Efficiency
张璐 believes that if the entire industry integrates AI, NVIDIA alone will struggle to meet demand; TPUs, CPUs and various specialized architectures will all have room. TPU adoption also reflects supply shortages and second-source considerations—no large company wants to place its entire future on a single vendor.
TPU’s strongest use case remains Google’s full-stack internal optimization. 张璐 recalls that Google’s internal training cost may be only about one-third of ChatGPT’s, but external customers may not reach the same level; NVIDIA has spent years serving many types of third-party customers, and CUDA and its system software remain more mature in general-purpose adaptation.
肖志斌 considers Google a “real, tangible threat.” It has iterated on TPU continuously since 2017 and may even be stronger than NVIDIA in interconnect, vertical power delivery, embedding and sparse acceleration, while also owning Gemini, YouTube and application distribution. As AI lowers the barrier to operator and system optimization, TPU will become increasingly usable externally.
NVIDIA still has 2 hard advantages: the execution required to launch multiple chips a year, and the supply-chain relationships 黄仁勋 has cultivated over many years. Even if AMD or Google wins orders, 肖志斌 believes NVIDIA still commands substantial TSMC and CoWoS capacity; trust and priority access to capacity are not things software can replicate quickly.
12. Edge, Private Deployment and Robotics Are the Flanks NVIDIA Has Not Fully Captured
肖志斌 believes robotics AI chips have not yet converged. When he assessed the direction 2 years ago, chips accounted for only 7%-8% of robot BOM and were not the bottleneck; the relevant models were also immature. 泓君 suggests that non-data-center AI could become an entry point that feeds later demand into data centers.
NVIDIA has reacted quickly to private deployment at the edge, launching Jetson small boxes and workstations capable of running larger models. 肖志斌 sees this as a defensive move: “黄仁勋 is quite concerned about private deployment at the edge.”
张璐 adds that traditional regulated industries care about data privacy and therefore favor private deployment. In discussions with Broadcom, she saw companies such as Qualcomm betting on edge AI. As inference takes a larger share, AMD and CPU solutions may also benefit; the market will ultimately form multi-architecture combinations by use case.
She identifies a nontechnical risk: the larger NVIDIA’s market cap becomes, the more management must preserve short-term revenue and capital-market expectations, potentially changing the resource balance between long-term R&D and near-term results. “Capital both celebrates it and holds it hostage.”
13. Samsung and Intel Can Serve as Second Sources, but Commercial Boundaries Are Harder Than Technical Validation
Christina asks whether Samsung and Intel can share NVIDIA’s foundry and packaging load. 肖志斌 says Intel’s EMIB packaging technology is strong in its own right, and the industry is testing whether TSMC-produced dies can be handed to Intel for packaging. The difficulty is that TSMC may not want to release the wafers, while Intel may require customers to adopt its own FPGA, creating a commercial paradox.
Samsung’s yields may be lower than TSMC’s, but it still has value as a second source, especially because NVIDIA releases 7 chips a year and different products can be matched to different suppliers’ strengths.
肖志斌 says Blackwell’s use of Samsung could also reflect capacity constraints or TSMC order-schedule issues, but stresses that this is only a guess: “Maybe that’s the case; I can only guess because I don’t know.” He also emphasizes that NVIDIA has always evaluated each supplier’s capabilities and did not arrive at the decision on a whim.
14. Coding Agents Weaken CUDA’s Kernel Moat, but Still Cannot Replicate Full-Stack Know-How
Mark acknowledges that coding agents can generate large amounts of code, but whether they can consistently write the highest-performance kernels remains unproven. Even if some CUDA capabilities can be replicated, NVIDIA’s moat has expanded into hardware, systems, toolchains and infrastructure; it is no longer a chip-only problem.
Mark relays feedback from engineers at large technology companies: AI-generated kernels can already reach more than 90% of the performance of hand-optimized kernels, so CUDA’s kernel-level moat has genuinely weakened. But the hardware data, failure experience and know-how accumulated at the system layer remain outside the grasp of general-purpose coding agents; that proprietary data will become a new moat.
张璐 adds the developer community dimension. Competitors have spent years trying to build CUDA-like ecosystems and recruit large software companies, but have still failed to truly replicate it. Through the Inception Program, NVIDIA has expanded its startup ecosystem from hundreds of companies to more than 20,000, cultivating potential customers and developer inertia in advance.
The discussion also reaches IP risk. When companies hand internal know-how to external systems such as Claude or Codex, they need to consider whether the models continue learning. Game studios are especially cautious; core data and low-level permissions require private deployment. 张璐 also says that, in conversations with Google teams, she found them unwilling to hand complete low-level control to an Agent.
15. US Data Centers Do Not Lack High-Voltage Power; They Lack Power Delivered to the Rack on Time
Alex’s most direct judgment is that data-center construction is moving fast, but the ultimate bottleneck remains “land and power.” US grid interconnection is nearly “bone dry”: new projects cannot secure more than 10MW of ready power in time, so roughly 90% of new data centers are turning to behind-the-meter generation.
The new model is to deploy gas generators at brownfield sites with natural-gas pipelines, generating power on site, stepping down the voltage and building the data center there. When grid access is blocked, operators may use diesel generators or natural-gas units directly. Many of the concerns hyperscalers had in the past are now largely ignored, and containerized construction is becoming standard.
The US is not short of generation capacity or 330kV high-voltage transmission; the problem is distribution. Power must be converted into the 400V usable by data centers, and now increasingly into 800V, while also passing through substations, stability studies and regulatory approvals. Alex says the grid is run by a traditional oil-and-gas system: “Their actions simply aren’t as fast as Silicon Valley.”
A single natural-gas plant typically produces 300-500MW, while a nuclear plant produces roughly 2-4GW; both remain insufficient for data-center clusters requiring hundreds of megawatts or multiple gigawatts. Alex says major companies in both China and the US are beginning to contract entire nuclear-plant outputs: “Don’t sell it to the grid—give it all to me.”
16. Modular AI Factories Cut Go-Live Times to 6-9 Months, but Component Shortages Extend Through 2027
A traditional greenfield project takes 18-20 months from land and concrete to service. Alex estimates that land, concrete, power, data-center white space, raised floors, water, electricity and fiber take about 4 months, followed by another 2-4 months to install racks and servers. A 40-foot container preloaded with racks, CDU, fiber, HVAC and UPS can cut the total cycle to 6-9 months.
This is how the “AI factory” becomes a supply-chain product. NVIDIA sells more than GPUs or a single server configuration: it provides an annual roadmap so power, racks, liquid cooling and ODMs can prepare on a common cadence, with racking and stacking standardized as far as possible.
GMI Cloud has locked in capacity through 2027, but Alex emphasizes: “At least we know capacity exists, but the price is impossible to determine.” Memory prices have risen 100%-200% since last year; HBM is squeezing DDR capacity, DDR is starting to go short, and SSDs are also becoming scarce. CX7 is shifting to BlueField, where the relevant solutions are primarily NVIDIA-based and lead times continue to lengthen.
The warning has also spread to CX7, switch gear, Intel CPUs and CDU liquid-cooling solutions. Based on supply-chain conversations, Alex expects no meaningful relief from these shortages at least through the end of 2027; the $1T order figure is therefore a simultaneous stress test of every physical component and delivery node.
17. GPU Cloud Really Sells Reliability, and Useful Life May Outlast Wall Street’s Depreciation Models
Alex reduces GPU-cloud operations to 1 word: reliability. New customers first demand “cards available and able to go live,” and only then stability. A system contains more than 200,000 unique parts and connects to thousands of devices; any nonzero failure rate will surface repeatedly at scale.
Troubleshooting spans hardware, optical modules, switches, firmware, K8s and customer code. Many model teams have top researchers but are not infrastructure teams, and the speed of GPU iteration leaves them without experience on new platforms. Cloud providers must identify root causes quickly, coordinate spare parts and restore service, ultimately delivering against the SLA.
Model services and token-cost reductions work only after operations are stable. Alex mentions optimizing inference through P/D, EP and clusterization; once a cloud provider reaches sufficient scale, scheduling and kernel optimization create room to reduce token cost.
On depreciation, he distinguishes the “hedge-fund answer: 5 years” from technical reality. AWS still struggles to make A100s or V100s available for rent, while V100s launched in 2017 or 2018 still see heavy utilization 7-8 years later. Excess demand means the economic life of GPUs may be materially longer than the 5-year depreciation period used by capital markets.
18. Neo Clouds Target the Gaps Left by Hyperscalers with Bare-Metal Efficiency and Product-Layer Services
Alex says traditional hyperscalers are fundamentally CPU and storage clouds, and their standard VM model gives up roughly 10% of compute power. That was tolerable when a CPU server cost only $20K-$30K, but a GB300 system is like “a house worth several million dollars,” and a GPU cloud cannot accept the same waste.
Neo clouds therefore prefer to use K8s to manage clusters while giving customers the full efficiency of bare metal. In competing with Google Cloud and Azure, the differentiation lies in cluster architecture designed for expensive GPUs, delivery speed and specialized operations.
Alex says GMI Cloud is 1 of NVIDIA’s 7 global Reference Architecture NCPs and can receive first-batch GPUs in sync with hyperscalers. It was the first company in Asia to receive a GB300 cluster, has built a 10,000-GPU liquid-cooled cluster, manages about 9 data centers, has another 3-4 under construction, and plans to bring its first scaled Vera Rubin deployments online in Q4 this year.
He places the next layer of differentiation in products: deployments across Asia and the US to meet data-security and low-latency requirements, followed by K8s, model services, workflow, studio and kernel optimization. The goal is not simply to rent cards, but to be a “one-stop shop” where enterprises and creators can call different models on reliable GPUs while continuously lowering token costs.