World Models—Do They Have a Scaling Law Too?
Summary
- World models may have a scaling law, but 庄明浩 defines it as a higher-friction version than LLMs: more expensive, slower and more dependent on real-world data. LLMs can absorb text from the internet at low cost, while physical AI needs vehicles, robots, sensors, test sites, safety oversight and customer deployments; feedback is slower and the long tail is more complex. The flywheel gets heavier as it grows, but it may also accelerate once it reaches a critical speed.
- The industry is closer to the GPT-2 phase than to the ChatGPT moment. Models can already generate worlds, predict simple states and control vehicles or robots in specific tasks, but reliable causal understanding in an open world remains distant; 庄明浩’s view is that the direction is clear and capital is committed, while the emergence threshold remains out of sight.
- A world model is not a “more realistic video model,” but the convergence of rendering, simulation and planning. An ideal model could render a cup from any angle, simulate its trajectory after being knocked over and plan for a robotic arm to pick it up safely; the real question is, “If I take an action in this world, what happens next?”
- Autonomous driving is the most likely first proving ground for physical AI’s scaling law because production vehicles are already terminals for continuously collecting real-world data. Cars operate mostly on relatively two-dimensional roads, with clearer safety requirements, data scale and business models than robotics; but data moats mean world models may not produce an LLM-style winner-takes-all market, with vertical models in autonomous driving, robotics, industry and gaming coexisting for years.
- Momenta’s core asset is not a single autonomous-driving algorithm, but a loop of “more cars—more data—better models—more OEM partnerships.” It had more than 680,000 installed vehicles as of end-2025; the episode puts its current production scale above 900,000 units and its market share at probably more than 65%, with 24 OEM partners, including 9 of the world’s top 10 automakers; the diversity of data across vehicle types, price bands, regions and driving styles is difficult for any single automaker to replicate.
- Momenta’s commercialization and engineering execution show that its flywheel has a base to get moving, but spillover from autonomous driving into a general-purpose world model remains a distant possibility that still needs to be proven. Citing the prospectus, 庄明浩 says revenue rose from more than RMB700M in 2023 to more than RMB2B in 2025, while gross margin climbed from 17.5% to 71.6%; the company has delivered more than 100 production vehicle models and cut the time needed to deploy 100,000 vehicles from 24 months in its early phase to less than 40 days at best. “Physical AI must have a cash-flow business to support it” is the basis of its “ticket theory.”
- The metric that matters is not whose demo looks most impressive, but who can keep the flywheel turning across safety, cost, performance and commercialization. R7 is built around an end-to-end foundation model, reinforcement learning and a world model, advanced through pretraining, closed-loop simulation and reinforcement learning; one risk is conflating a mature autonomous-driving business with cross-domain generalization that has yet to be validated.
Deep dive
1. Three Signals Show That World Models Have Become the Common Language of Fundraising and Commercial Narratives
影视飓风’s experiment for Tmall’s 618 shopping festival posed the question bluntly: “If you flip a coin in AI, is the probability 50%?” The team ran more than 100 tests, with heads accounting for more than 70%; 庄明浩 argues that this exposes not a random-number problem, but how training-data distributions, post-training preferences and reinforcement-learning biases enter generated worlds.
In April, BAI Fund partner 王天凡 observed that among mainstream funds, “there are as many world models in mainstream fundraising as there are partners.” In June, he said, “I receive at least one world-model BP every day—it feels like being back in Series B.” The heat has spread from a research topic into a standard keyword for startup pitches.
The third signal comes from Momenta, which has passed the Hong Kong Stock Exchange’s listing hearing: a self-driving company with production revenue has begun explaining itself through real-world data, reinforcement learning and world models. That turns world models from a lab-level technical thesis into a commercial example whose data, revenue and delivery capabilities can be examined.
2. Language-Model Limits and Massive Capital Are Driving the “Next Main Line”
庄明浩 describes the state of language models over the past 6 months as “unstoppable”: writing, coding, summarization, reasoning, tool use and autonomous evolution continue to advance. He estimates that model-related fundraising accounts for more than 60%—and possibly 70%—of the US primary market. As leading assets grow larger, early-stage capital has little choice but to search for opportunities beyond language.
Menlo Ventures counts roughly 63 New Labs within its coverage universe, with cumulative funding potentially reaching several tens of billions of dollars. The fund’s latest early-stage and growth vehicles raised $3B in aggregate, versus about $1.3B in the previous cycle. New Labs routinely raise tens of millions, hundreds of millions or even more than $1B, in turn forcing funds to increase their own scale.
伊利亚, Mira, 杨立昆 and 李飞飞 are among the star founders and researchers in this New Lab cycle. Their companies and teams span autonomous evolution, physical AI, humanoids and embodied intelligence, world models, AI for Science and model infrastructure. 庄明浩 repeatedly uses the word “thesis” because the industry has not converged on a shared architecture; capital is still funding a view on the path, not a validated standard answer.
3. World Models Need to Understand the World Itself, Not Merely How Humans Describe It
李飞飞’s analogy is that an LLM is “a language master in the dark”: articulate and well-read, but without direct experience or grounding. “The world is not made of words.” A model may know that a cup breaks simply because text says so, rather than because it has continuously observed light, geometry, gravity, friction and the flow of liquids.
What world models compress is space-time and physical law: a cup follows different trajectories on a tabletop, concrete or ice; a vehicle changing lanes alters the reactions of nearby drivers and pedestrians; a robotic arm peeling a raw egg versus a boiled egg must understand force, material and deformation. None of these is merely a language problem.
The core task, therefore, is not to generate a video that “looks plausible,” but to answer: “If I take an action in this world, what happens next?” Attractive imagery can tolerate fabrication. Braking a vehicle, grasping an object or predicting a collision requires the model to be accountable for the consequences of an action.
4. Renderers, Simulators and Planners Are Converging on the Same Foundational Capability
李飞飞 divides current approaches into 3 categories. Renderers output pixels and optimize for visual realism; simulators output environmental states and must get geometry, dynamics, vehicle trajectories and collision relationships right; planners output actions, determining how a robot grasps or how a vehicle turns, accelerates or decelerates.
庄明浩 draws the boundary with a rainy-night highway obstacle-avoidance scenario. A video model can reproduce camera shake, reflections on wet roads and the look of headlights, but may not know when to brake, whether the vehicle will be rear-ended, whether turning left will intrude into the adjacent lane or what risks each strategy carries. “This is not a rendering problem.”
An ideal model would understand 3 projections of a cup: render its lighting and appearance from any angle, simulate its motion after being knocked over and plan for a robotic arm to pick it up without breaking it. The 3 capabilities share geometry, physics, dynamics and object-interaction rules at the foundation, making the convergence of the 3 lines more relevant to the end state than simply improving video sharpness.
5. Behind the Capital Frenzy Is Real Anxiety, Not Just a New Buzzword
World models have become an all-purpose container much like the “metaverse” before them, but 庄明浩 is not rushing to call a bubble. He is more interested in the question behind the heat: if “the right prediction task + data + model + compute” produced a scaling law for LLMs, can predicting the next physical state follow the same path?
The episode cites a CB Insights report for Q1 2026 showing physical AI/world models at roughly 11%, the largest category among fragmented AI subsegments. The source first describes this as a share of “project investment,” then refers to nearly 11% of “funding,” without clarifying the denominator. China’s IT Juzi report counts more than 30 related startups, over 85% of which were founded after 2023, with cumulative funding in the tens of billions of yuan.
DeepMind, Nvidia, video-model companies, autonomous-driving firms, robotics companies and AI 3D startups are entering the same battlefield from different directions. 庄明浩 extends his earlier view that multimodal AI is moving from vertical battlefields toward a unified one: that unified battlefield now includes autonomous driving and embodied intelligence, not just video models.
6. The Cost and Density of Physical Data Mean Scaling Law Will Not Transfer Smoothly
LLMs can obtain webpages, books, encyclopedias and code relatively easily. World-model data requires vehicles to actually drive, robots to actually work and industrial equipment to actually run. Every data point carries hardware, facilities, labor, safety and compliance costs; “crawling a webpage and doing this cannot possibly be in the same order of magnitude.”
Physical data may look abundant while carrying less semantic density. The sentence “The driver sees a child suddenly run out, brakes hard and swerves left” compresses the key facts. Real-world collection requires several seconds of video, camera and radar feeds, vehicle trajectories and labels, while the truly valuable event may occur in only a single instant.
Training a world model at an LLM-like scale may therefore require not just somewhat more data, but several orders of magnitude more. Video volume alone is insufficient: teams must filter for high-quality samples and accumulate enough of the long tail to cover the tighter safety boundaries of physical systems.
Scaling law thus becomes a supply-and-closed-loop problem rather than a simple “more is better” equation. Can a system continuously secure real-world data? Is the sample distribution diverse? Is the long tail sufficiently covered? Which samples are high quality? Can the outputs feed back into the training environment? If any link breaks, scale may not automatically translate into capability.
7. Autonomous Driving Is Closest to a Closed Loop, but No One Can See the Emergence Threshold
Compared with robots, production vehicles already drive every day across different roads, cities, weather conditions and traffic environments. Each vehicle is a data-collection point, and users’ driving itself becomes a source of content. Autonomous driving therefore has the best chance of validating world models first.
庄明浩 is explicit that “as of this point in time, emergence has not yet appeared.” Existing models can generate worlds, predict simple changes and control vehicles or robots in specific tasks, but they still cannot reliably understand a complete causal chain in an open world. No one can yet observe how large the data set must become to cross the threshold.
Forced into a GPT analogy, he says world models are “roughly at GPT-2”: not yet at the point where GPT-3 had made the light visible, and even farther from the GPT-3.5 phase that produced ChatGPT. Capital intensity, excitement around demos and pockets of commercialization are already present, but they do not mean general-purpose capability has arrived.
8. Three Star Approaches Are Still Competing to Define the World’s Basic Representation
杨立昆’s thesis is to avoid generating pixels directly and instead predict state changes in latent space. 庄明浩 characterizes the benefits as relatively high compute efficiency and stronger generalization, with poorer interpretability as the trade-off. The bet is on abstract representation rather than the visual output itself.
DeepMind is pursuing a multimodal-plus-evolution route, adding vision, spatial perception and reinforcement learning on top of an LLM architecture. Genie 3’s demonstrations still begin with video generation and operability. 庄明浩 sees the approach as relatively pragmatic and already showing some results.
李飞飞’s World Labs is betting on 3D spatial representation: generate the world itself first and treat spatial and physical representations as the base layer. The 3 frontier scientists’ theses have clearly diverged, but the episode declines to declare a winner at this stage.
9. Momenta’s Data Flywheel Rests on Production Scale and Cross-OEM Diversity
Momenta’s loop is straightforward: production solutions generate revenue and real driving data; the data trains and improves the models; better models feed back into production programs and L4 efforts, attracting more OEM partnerships. “More cars bring more data, and more data brings better models.” The hard part is securing the vehicles and customers needed to start the flywheel.
As of end-2025, Momenta had more than 680,000 installed vehicles. The episode puts its current production scale above 900,000 units, with market share probably above 65%. The company works with 24 automakers globally, including 9 of the world’s top 10, and several of those automakers are also strategic investors.
Unlike Tesla or a single Chinese automaker building in-house, Momenta’s data comes from different manufacturers, vehicle models, price bands, regions and driving styles. 庄明浩 argues that the narrower the training distribution, the harder generalization becomes; cross-brand diversity may be a more important moat than raw data volume alone.
This remains a classic chicken-and-egg problem: no vehicles means no data, no data makes model improvement difficult, and without a better model it is hard to win OEM customers. Momenta’s current advantage is that it has crossed the initial commercial-deployment threshold, not that the flywheel can never be caught.
10. Cash Flow, Engineering Delivery and R7 Form Physical AI’s “Ticket”
Citing the prospectus, 庄明浩 says Momenta’s revenue rose from more than RMB700M in 2023 to more than RMB2B in 2025, while gross margin increased from 17.5% to 71.6%. Licensing revenue is rising, shifting the mix from project-based work toward software licensing that increasingly resembles SaaS. The company remains loss-making, but the loss is narrowing.
CEO 曹旭东 calls this the “ticket theory”: physical AI needs a cash-flow business to support it because vehicle testing, deployment, compliance, maintenance and OEM integration cannot be distributed globally and instantly like an app. If moving from GPT-2 to GPT-3.5 takes a long cycle, a pure R&D company without cash flow may be eliminated by the financing environment first.
Engineering execution matters just as much as paper metrics. Momenta has delivered more than 100 production vehicle models, adapting to different vehicle platforms, sensor configurations, quality and safety standards, costs and delivery schedules. Its first 100,000 production vehicles took 24 months; today, the fastest deployments can deliver 100,000 vehicles in less than 40 days, indicating that deployment is becoming platformized.
Technically, the company moved from a modular deep-learning system in 2022 to an end-to-end foundation model in 2024-2025. R6 introduced reinforcement learning across the full process in 2025, while R7 explicitly combines “an end-to-end foundation model + reinforcement learning + a world model,” with reward objectives covering safety, comfort and efficiency.
11. R7 Breaks the World Model into 3 Layers, but the ChatGPT Moment Will Still Arrive Gradually
The first R7 layer pretrains on massive volumes of real driving data, compressing physical laws, common sense and causal relationships into the model. The second builds a closed-loop simulator, allowing the model to reason through how the world changes after an action and evaluate rare, difficult-to-reproduce long-tail scenarios that occur on real roads.
The third layer applies reinforcement learning in a virtual environment that closely approximates reality. The model no longer merely imitates human drivers, but uses trial and error to find safer and more efficient decisions. “Imitating humans always has a limit.” Whether the team can design effective rewards and elicit better behavior will determine whether the architecture can truly take off.
The likely rollout sequence is autonomous driving first, followed by gaming, virtual worlds and interactive film content, where the tolerance for error is higher, and only later robotics, industrial simulation, digital twins and AI for Science. Robots face full 3D space, deformable objects, hand-eye coordination and material deformation, making them far harder than vehicles operating on two-dimensional roads.
World models may also avoid an LLM-style concentration among a handful of winners. Cars, robots, warehouses, industrial systems, surgical robots and games have different sensors, action spaces, risk boundaries and commercial constraints. The underlying representation may be shared, but vertical data moats will persist, leaving Momenta’s first opportunity in autonomous driving rather than control of the entire physical world.
A ChatGPT moment requires direct mass-market perception, a very low product barrier and sufficient generality to trigger a broad supply-chain boom. What world models currently offer is mostly spectacular video, robot marathons and demonstrations across different scenarios. “The physical world does not allow you to be approximately right.” The real test is therefore no longer whose demo is strongest, but who can balance safety, cost, performance and commercialization over the long term.