Vol.71 A Quick Chat About GPT-5 and U.S. Power—with 刘一鸣/101 Weekly
Summary
- GPT-5 did not prove that the data wall has disappeared; it showed that reinforcement learning can extend the scaling law through new routes. 庄明浩 believes Ilya was right about the original-data bottleneck he identified in September 2024; what has changed is the combination of synthetic data, pretraining reuse, and post-training—a self-lifting loop in which “the left foot steps on the right and flies upward.” Mathematics, coding, and medicine are precisely the fields where reinforcement learning can deliver clear gains.
- Compute demand will continue shifting from training to inference, and falling unit costs will not reduce total demand. 庄明浩’s order-of-magnitude estimate is that token costs fall by roughly an order of magnitude every 12 months, while token consumption and required compute also rise exponentially over the same period—“it evens out.” Inference tokens and compute are not simply linearly related; time and per-use costs also matter, so the problem cannot be solved with a single equation.
- GPT-5’s launch fell short of expectations, but product engineering is still filling gaps on the To C front. Reducing hallucinations, improving routing, removing sycophancy, and implementing finer-grained engineering changes are all improving the product experience; better tool use, longer tasks, and higher success rates point to a more reliable agent. 庄明浩’s view is that “the main battlefield is still To C.”
- OpenAI has to push harder into To B because consumer subscriptions are unlikely to cover expanding CapEx on their own. ChatGPT’s weekly active users rose from roughly 300M–400M at the start of the year to 700M, but even with further growth, weekly actives multiplied by conversion and then by monthly ARPU of $20 or $200 still produces an estimable number; enterprise implementation and government contracts offer larger marginal gains, while competition from Anthropic and cloud providers must also be confronted directly.
- The broad scaling law still holds, but capability growth has shifted from “feeding in more data” to combining pretraining, post-training, and reinforcement learning. 庄明浩 believes this R&D path was still working at least through Q3 2024; GPT-5’s gains in tool use, task length, and success rates fit the L3 agent phase. As for the next milestone and how L4 will be achieved, “we outsiders may simply have no way of knowing.”
- The re-rating of U.S. power assets is being driven by data-center CapEx reaching a critical point for the power system, not by GPT-5’s single release. Several leading companies are now spending roughly $100B per quarter; 庄明浩 relayed a Morgan Stanley report saying annual investment could reach $3T by 2029 or thereabouts, leading him to speculate that “the entire U.S. power ecosystem could be rewritten.”
- The opportunity in power expansion extends beyond generation equipment to the full chain of production, transmission, distribution, and auxiliary infrastructure. 刘一鸣 is watching GEV’s steam gas turbines, steam turbines, SMRs, hydropower, and related businesses, noting that its order book runs through 2029; both speakers drew a clear line around their judgment—they are extrapolating from AI demand and CapEx, and do not yet know whether the data has reached a critical point or exactly what needs to change.
Deep dive
1. Reinforcement Learning Bypasses the Original-Data Bottleneck Without Disproving the “Wall”
刘一鸣 first asked whether the data bottleneck Ilya identified in September 2024 still holds. 庄明浩’s answer was unequivocal: “Ilya was not wrong.” Original data in the public world, along with the marginal efficiency of reprocessing and relabeling it, has indeed deteriorated sharply.
The real surprise over the past year has been the growing weight of reinforcement learning. It is no longer merely a post-training tool for improving performance; trained models can now generate synthetic data that feeds back into pretraining, creating a self-reinforcing loop.
庄明浩 described the mechanism as “stepping on your left foot with your right and flying upward”—a self-lifting loop. Once pretraining, synthetic data, and reinforcement learning can reinforce one another, capability gains are no longer fully constrained by additional internet-scale text.
GPT-5’s focus on mathematics, coding, and medicine was not accidental. 庄明浩 believes all 3 areas offer clear pathways for reinforcement learning to improve model performance. The leading-model releases and high scores from DeepSeek, Qwen, Llama, Anthropic, and OpenAI over the past year broadly fit this trajectory, while Google “may be somewhat different.”
2. Inference Demand Eats the Cost-Cutting Dividend as GPT-5 Keeps Moving Toward Agents
Moving from base models to reasoning models, then adding model scale, complex data, and synthetic data, ultimately raises token consumption across training, inference, and applications. The shift in compute demand toward inference should therefore continue.
庄明浩 used the logic of “Andy and Bill’s law” to explain why lower costs do not mean lower total consumption: token costs fall by roughly an order of magnitude every 12 months, while consumption and compute demand also grow exponentially over roughly the same period. “It evens out.”
刘一鸣 asked whether a 10x increase in tokens implies a 10x increase in compute. 庄明浩 refused to force a simple answer: “It may involve more than just compute; it may also involve time.” Compute, time, and per-use costs make this “data that cannot be solved with a single equation.”
3. GPT-5 Delivered No Surprise, but OpenAI’s Commercial Pressure Is Becoming More Concrete
刘一鸣 viewed GPT-5’s overall launch as “not particularly impressive.” 庄明浩 countered that product improvements had not disappeared: reducing hallucinations, improving routing, removing sycophancy, and adding more fine-grained engineering work are all improving the user experience.
庄明浩 believes the main line for OpenAI and Sam remains To C. The problem is that ChatGPT’s weekly active users have already grown from roughly 300M–400M at the start of the year to 700M, making it difficult to demand another exponential increase from here.
Consumer revenue can be roughly decomposed into user count, conversion rate, and monthly ARPU of $20 or $200. Even if OpenAI reaches the previously discussed 1B daily active users, the figure may still be “not enough” relative to compute, CapEx, and other costs, leaving To B as the only source of faster incremental growth.
Enterprise implementation teams and large government contracts are manifestations of this strategy, but the competitive environment is tougher. 刘一鸣 believes Anthropic “has hit OpenAI very hard” in enterprise services and coding, while cloud providers hold important positions in the To B market, forcing OpenAI to invest more.
4. The Broad Scaling Law Is Still Advancing; What Comes After Agents Remains Unknown
庄明浩 defines today’s scaling law more broadly: it no longer means only feeding more data into pretraining, but also includes post-training, reinforcement learning, and related methods. At least through Q3 2024, he believes leading companies still had relatively clear stage goals, milestones, and implementation paths.
庄明浩 recalled OpenAI’s earlier description of the roadmap as moving from chatbot to reasoning and then to agents. After agents may come an “organizer” that assembles different capabilities, followed eventually by innovation—though he added, “if I remember correctly.” GPT-5’s stronger tool use, longer task completion, and higher success rates fit the requirements of the L3 agent phase.
On the path from agents to higher stages, 庄明浩 remains uncertain. Leading companies “may” already be experimenting internally, but outsiders do not know exactly what they are testing or where they are directing their efforts, and cannot tell how L4 will be achieved.
5. Data-Center CapEx Has Pushed the U.S. Power System to a Critical Point
刘一鸣 noted that power stocks such as GEV had at one point risen even more than Nvidia. 庄明浩 looked back to Q3 2024, when infrastructure and energy “should have ranked first” among S&P 500 subsectors, even ahead of IT and technology. The groups retreated in Q4, then returned to the discussion in January and February after the DeepSeek event.
The turning point came with the Q2 earnings reports from several leading companies. Application performance, user expansion, revenue and profits, and cloud-revenue growth all came in well, rapidly reversing market sentiment from bearish to bullish. Combined with 2–3 years of sustained data-center CapEx, 庄明浩 inferred that the buildout may have reached a critical point, beginning to affect the stability and capacity of the power system and even residential users. The commercialization of the U.S. power market and its link to prices further pushed the market into a new state.
Several leading companies are now spending roughly $100B per quarter. 庄明浩 relayed a Morgan Stanley report saying annual investment could reach $3T by 2029 or thereabouts. If this investment trend were to multiply several more times, “the entire U.S. power ecosystem could be rewritten.”
The data-center construction curve is rising rapidly while office construction has nearly collapsed; the 2 are approaching a crossover and could eventually see data centers surpass offices. If the trend persists for another 3 years, nuclear plants and other power facilities could bring some Rust Belt cities that have faded from public discussion back into view.
6. The Opportunity Runs Through the Entire Grid-Equipment Chain, Not Just Generation
刘一鸣 stressed that U.S. power expansion resembles a new round of “big infrastructure.” New power plants solve only generation; transmission, distribution, transformers, and other auxiliary infrastructure face bottlenecks as well, creating opportunities for another group of listed companies.
Economic conditions and grid infrastructure vary widely across U.S. states, and some regions may need to be redesigned from the ground up. 庄明浩’s reverse framing was: “If your foundation is weak, the room for improvement may be greater.”
刘一鸣 described GEV as covering steam gas turbines, steam turbines, natural-gas generation, SMRs, and hydropower—essentially touching multiple power-related themes—and said its order book extends through 2029. But neither speaker presented the reasoning as a precise forecast: “We are not energy experts.” Whether the system has reached a specific critical point, and what exactly needs to change, still requires validation with specialized data.
At the end of the episode, 刘一鸣 said GPT-5 had limited impact on compute and other areas, and Nvidia and related names showed little movement the following day. Some compute-leasing companies rose, apparently in response to the earnings reports of certain companies rather than GPT-5 itself.