170: With 陈哲’s “Embodied Quarterly 26Q2”: The World-Model Gale Keeps Blowing, and Those Who Don’t Want Labels
170: With 陈哲’s “Embodied Quarterly 26Q2”: The World-Model Gale Keeps Blowing, and Those Who Don’t Want Labels
Summary
- World models are moving from concept to product and paradigm. Nvidia’s Cosmos 3, released on June 1, is the quarter’s benchmark: the market’s first fully open-source Omni Model, using a Mixture of Transformers architecture to “stitch” autoregressive reasoning and diffusion generation into a unified model, with super, nano, and edge versions released at once for deployment. 陈哲’s view: “Rather than saying world models are overturning VLA, they are bringing new capabilities and ideas to existing SOTA models.” Pi 0.7’s connection to a lightweight world model and Generalist Gen One’s refusal to be categorized as either VLA or world model are both signs of convergence.
- Honor’s clean sweep of the Yizhuang Marathon podium was a preview of the big manufacturers’ entry. All 3 of its autonomous-navigation “Lightning” robots finished in roughly 50 minutes, versus Tiangong’s 2 hours and 40 minutes to win the remotely controlled race in 2025—a more than 3x improvement in a year. Honor kept the motors from overheating with race-specific motors and precision liquid cooling, spending “significantly more than the other 2 teams.” The host’s question was whether a body-maker’s ability to IPO this year could become a resource-sufficiency line; 陈哲’s answer was that Xiaomi, XPeng, Li Auto, and other major end-device manufacturers are already entering seriously, and humanoid robotics is “evolving from single-technology entrepreneurship into a system of systems engineering and systematic execution.”
- Figure’s 100-plus-hour livestream sorting 130,000 parcels marked the humanoid robot industry’s “zero-to-one” moment. At roughly 1 parcel every 3 seconds, flipping and flattening soft packages—a motion that was “basically unsolvable with the technology stack of 4 or 5 years ago”—the robots went beyond the pick-and-place limits of suction-cup systems. The teleoperation debate is beside the point: “I believe every humanoid robot goes through a long period of teleoperation before entering a deployment scenario like this,” while remote intervention is already the norm in robotaxis. 星动纪元’s comparable deployments with China Post and SF Express are “in no way behind” those of US companies.
- Dexterous hands are becoming the next competitive battleground, strengthening the case for direct drive. At ICRA Vienna, 五G released a second-generation direct-drive hand with 20 degrees of freedom, much better backdrivability and thermal management, and roughly half the volume of ShaPa’s hand—positioning it as “Unitree in dexterous hands.” “Whoever can provide reliable, stable, inexpensive dexterous hands faster will become the industry’s de facto standard faster.” Domestic major manufacturers’ move to Tesla’s tendon-drive route is more about organizational risk management: “If Tesla turns out to be right, and you try a new direction, who takes responsibility for that mistake?” Optimus Gen 3’s hand continues to iterate, while leading high-DOF hands ship only a few thousand units.
- Data is the next positioning battle. Collection paradigms have moved rapidly over 2 years from ALOHA real-robot teleoperation to UMI without a body, egocentric video, and full-body motion capture; “data we start collecting today could translate into breakthroughs in models 3 or 6 months from now.” Dexterous-hand data is highly dependent on hand geometry and sensor choices, making standardization by third-party data vendors difficult. Unlike lidar point clouds, which can run through a single pipeline, “dexterous hands and humanoid bodies will most likely be locked in a long-term contest and long-term coexistence.”
- Gen One was the quarter’s most discussed model advance. Generalist collected 500,000 hours of UMI-style, body-free data and trained an end-to-end model from scratch without a pretrained model, lifting success rates on complex long-horizon tasks from the low 60% range to 99% and claiming to have “found the key to scaling.” Peter Florence explicitly rejected both the VLA and world-model labels; the confidence behind that route reflects a US venture ecosystem that “rewards the first person to do something.”
- The endgame question is why embodied models would not belong to general-purpose model companies. “What reason do we really have to believe that, many years from now, the model for embodied AI will not be provided by Anthropic, OpenAI, or Google?” Cosmos 3’s 10K-GPU-scale training is already the ceiling for embodied-AI research compute, yet not an especially large budget for leading LLM companies. OpenAI’s high-profile robotics restart draws from the Sora team, whose internal positioning has always been world models, with dexterous hands likely to be a major investment area. 陈哲 expects embodied models to end up like language models: “high-quality open source plus top-tier closed source, with no room for other models to survive.”
- Unitree’s IPO approval could become the industry’s valuation anchor, while the next 6 months may be the last window for major manufacturers to enter. “If you do not start entering this market in 2026, by the time humanoid robots are broadly deployed in 2027 or 2028, you will have no position.” The China-specific paradox is that investors tolerate neither 10 years without commercialization nor a company that is not full-stack: “Every embodied-brain company is announcing that it is building a body.” Yet the companies with the largest market caps are precisely not full-stack: Nvidia is not, and Anthropic may not be either if it goes public.
Deep dive
1. Q2’s 5 Biggest Embodied-AI Developments: Breakthroughs in World Models and Dexterous Hands Bring Humanoids to the Masses
- 陈哲’s ranking: the Yizhuang humanoid marathon, with more than 100 units competing and Honor taking the title, signaling that “major manufacturers can build good humanoid robots”; Figure’s multi-day livestream of parcel sorting, which brought the industrial value of humanoids to a global mainstream audience for the first time; the breakout in dexterous hands and dexterous manipulation, led at ICRA by 五G’s new high-DOF hand and accompanied by multiple startups building dexterous-manipulation foundation models; Nvidia Cosmos 3, which “turned the world model from a concept into a real product and a benchmark paradigm,” and became the quarter’s hottest startup theme; and continued iteration on the VLA route through Pi 0.7 and Gen One, where “signs of fusion with world models are beginning to appear.”
- Industry terminology is also shifting quietly, with more people using “physical AI,” while world models are carrying forward their top-5 momentum from last quarter.
2. Last Quarter’s Outlook, Put to the Test: World Models Move from the Lab to Industrial Grade
- When the show discussed world models last quarter, they were still “early lab-grade prototypes.” Nvidia’s Dream Zero was a small internal GEAR Lab project built on the open-source “Wan” video model. Cosmos 3, released in June, is “more product-grade and pretrained at much larger scale,” launching super, nano, and edge versions in one shot for different deployment environments—models that “can be called and deployed.”
- Asked whether world models could produce progress beyond VLA, 陈哲 deliberately avoided a replacement narrative. The strengths and weaknesses of classic VLA are already fairly clear; world models are better at environment prediction and modeling. “Rather than saying world models are overturning or surpassing VLA, they are bringing new capabilities and ideas to existing SOTA models.” Leading VLA models are already connecting to or “stitching in” world-model capabilities through different methods.
3. Yizhuang Marathon: Honor Sweeps the Podium, with Liquid Cooling the Decisive Edge
- The 3 pre-race favorites were Unitree, Beijing Humanoid Robot Innovation Center—the home of defending champion Tiangong—and Honor, and the market had already anticipated Honor’s strength. Unitree made its first official appearance and posted 10 meters per second in pre-race testing. Honor’s robotics division was formed about 2 years ago, with a team of nearly 100 to 200 people; its investment in the race was “significantly greater than that of the other 2 teams.”
- Where did the money go? Honor customized the motors for the race and added a precision liquid-cooling system. The livestream showed other models forced to rest after extended high-speed running because their motors overheated, while Honor’s cooling system kept the motors at relatively low temperatures throughout. That was the key reason the “Lightning” robots could maintain high speed and consistency across the entire marathon.
- There was also a strategic tell in the race format: Honor sent 10 teams across the teleoperation and autonomous-navigation categories, while Beijing Humanoid Robot Innovation Center and Unitree each sent only 1 or 2. Having ample resources is itself a competitive advantage.
4. Performance Improves More Than 3x in a Year: The Marathon Is Humanoid Robotics’ F1
- The numbers are stark. In 2025, champion Tiangong finished the remotely controlled race in 2 hours and 40 minutes, with Songyan Dynamics second at 3 hours and 37 minutes. In 2026, all 3 of Honor’s autonomous-navigation entries finished in roughly 50 minutes, tightly bunched. 陈哲: “Performance improved more than 3x in roughly a year,” reflecting rapid progress in robot hardware, control, and autonomous navigation.
- To the longstanding question of what the race is actually good for, his analogy is worth keeping: “F1 racing was never designed to sell cars.” The marathon creates an extreme test environment that forces advances in whole-body control, emergency handling, long-duration failure-free operation, power, and thermal management. Once engineers understand the boundaries of a humanoid’s capabilities, mass production can draw on substantial system-design experience. The other implication: “If everyone seriously optimizes this problem, many issues can be solved with today’s technology stack.” The industry simply had not previously demanded the limit of locomotion performance.
5. The Major-Manufacturer Signal: From Single-Technology Startups to “Systematic Execution”
- The host raised an industry concern: if startups that excel at bodies cannot IPO and reach a resource-secure state this year, large end-device manufacturers such as Honor—with businesses in phones and cars—will enter with their own bodies in 2027. They have more experience in mass production, reliability, and consistency. 陈哲’s response was that the marathon proved “teams with high-end manufacturing experience and strong organizational capabilities can quickly produce highly competitive humanoid bodies with enough resources and talent density,” while their control and navigation algorithms are also no weak link.
- The broader read is that China has quite a few companies with Honor’s organizational, talent, and financial resources. Xiaomi, XPeng, Li Auto, and other EV and smartphone manufacturers have begun investing seriously. Humanoid robotics, as “a systems-engineering discipline combining high-end manufacturing, complex algorithms, and complex software systems,” is “evolving from the perspective of single-technology startups into a system of systems engineering and systematic execution.”
6. Figure’s Livestream: 130,000 Parcels, 1 Every 3 Seconds, a “Zero-to-One Breakthrough”
- Starting May 13, 3 Figure robots livestreamed for more than 100 hours, standing beside a conveyor belt to flip each parcel and place its barcode face-up. They sorted 130,000 parcels, or roughly 1 every 3 seconds, with human pace and success-rate comparisons on site. 陈哲 put it in the top 5 because, “both in terms of spectacle and actual industrial value, it was a zero-to-one breakthrough.” Parcel sorting requires a degree of generalization and general-purpose manipulation, and is “a setting that is extremely unsuitable for humans to work in for long periods.”
- The host’s question was practical: the robots never moved from their stations during the livestream, so are bipedal legs and humanoid arms redundant? 陈哲 agreed that in most cases 2 hands are all that is needed to pick up and flip a parcel, but real operations contain surprises—conveyors moving too fast, objects too light, or spherical items sliding to the floor, as happened once or twice on the livestream. “Only a humanoid robot, combined with the generalization capabilities of today’s embodied models, can potentially solve the long-tail problems you cannot enumerate.” The robots cannot yet handle these corner cases.
7. Why the Old Stack Could Not Solve It: Soft Parcels and “Flip and Flatten”
- Companies using machine vision and industrial arms were trying to solve the same problem 4 or 5 years ago, and China Post had already been working with multiple robotics companies. But there was no true industrial deployment, only a surplus of demos. The problem was that parcels are often soft, while QR codes can be covered or deformed. Humans use 2 hands to flip the package and flatten it so the code is visible. “Those 2 actions were basically unsolvable with the technology stack of 4 or 5 years ago.” The more general suction-cup approach at the time offered almost no 2-handed operation—only basic pick and place.
- That is why both Figure and 星动纪元 use dexterous hands, and why the logic is similar to Pi and Dyna’s clothes-folding demos. Manipulating deformable materials is “very well suited to today’s embodied models”: soft goods, textiles, and plastic bags cannot be modeled geometrically with simple rigid-object models, so the system needs generalization and adaptability. Standardized warehouses filled with rigid cartons have already been addressed by HAI ROBOTICS, Geekplus, and others. What remains is the nonstandard demand in e-commerce warehouses, with many product categories and highly flexible materials.
8. The Teleoperation Debate Is Not the Debate: Remote Intervention Will Be the Norm
- A “head-support” motion during the livestream prompted questions about teleoperation; Figure said it was completed autonomously by the Helix 02 model. 陈哲 cut straight through the issue: “Teleoperation is fundamentally not the point of contention here. I believe every humanoid robot goes through a long period of teleoperation to collect data and correct actions before entering a deployment scenario like this. Even if humans need to intervene for part of the time, that does not diminish the significance of the deployment.”
- The longer-term model will resemble autonomous driving, with a human operator monitoring several robots from a control room. “Remote intervention is already a market norm for robotaxi systems. When we look at Waymo, we measure takeover efficiency and how many vehicles 1 person can monitor.” The boundary is the home, where privacy raises the bar for autonomy. “That is why humanoids entering the home will be more difficult than entering purely industrial settings, both from a privacy perspective and from a data-distribution perspective.” Historically successful B2B robots operate at high deployment volumes, high throughput, and in consistent environments. Some industrial humanoid deployments are simply doing what dedicated arms could already do; Figure chose a setting the old solutions had not actually solved.
9. China Is Not Behind: 星动纪元’s China Post and SF Express Tests, and a Second Curve for Legacy To-B Companies
- 陈哲’s information is that 星动纪元 has conducted extended testing and training in parcel sorting not only with China Post but also SF Express. It “has already achieved fully autonomous parcel splitting, flipping, scanning, and a series of related processes” with humanoid robots. “Chinese companies’ progress in deploying humanoids in logistics and industrial settings is in no way behind that of US companies.” It is a strong PMF: bimanual humanoids, embodied models with generalization, and dexterous hands. There should be no shortage of teams with all 3 capabilities.
- Could established logistics specialists such as HAI ROBOTICS and XYZ enter the market? 陈哲 returned to the underlying logic of To B: “Very few technologies are inherently monopolistic. More important is whether your product can be embedded in a larger operating system.” Industry knowledge and integration with customer systems are the real barriers. Most of the previous wave of To B robotics companies are already at the IPO or pre-IPO stage. Once they actually go public, “the resources and capabilities they can access may be far greater than those of today’s private-market companies.” Pudu is already a leader in commercial robotics, and warehouses and properties still employ large numbers of workers; tasks untouched by simple automation remain a broad extension opportunity.
10. Four Data Paradigm Jumps in 2 Years: ALOHA → UMI → Egocentric → Full-Body Motion Capture
- 陈哲’s methodological point is that data is the overall bottleneck in embodied intelligence, and “each model-paradigm iteration is really a change in the data paradigm.” Watching what data-collection companies are doing can therefore provide an early read on where the industry is headed. The progression: ALOHA real-robot teleoperation 2 years ago, including Tony Zhao and 子鹏’s 2023 work with 4-arm isomorphic cloning; UMI last year, using a handheld 2-finger gripper without a body and no real robot; and egocentric first-person video by the end of last year, captured with a head-mounted camera, with or without gloves, where “diversity exploded.”
- After Nvidia open-sourced its full-body motion-capture work, phonetically “Sonic,” at the end of last year, human movement and dance could be transferred rapidly to humanoid bodies. The algorithmic side relied on that open-source stack; the hardware side on Unitree’s mature body; and the open-source market had many people optimizing parameters around Unitree. This year has brought a wave of full-body control-data collection and loco-manipulation experiments, as well as new specialist companies such as 李宏洋’s 原测未来. The key line: “Data we start collecting today could translate into breakthroughs in models 3 or 6 months from now.” Locomotion control was once treated as a cerebellar function separate from the brain; now there is an opportunity to train movement and manipulation in a more unified architecture, increasing the value of teams with strong control capabilities.
11. ICRA Vienna: 五G’s Second-Generation Hand Stole the Show
- A large number of Chinese companies exhibited at this year’s ICRA, with dexterous hands the center of attention. 五G released a new direct-drive hand; 希诺, 临界点, and others launched high-DOF products; and 星动纪元 released a 21-DOF flagship hand, while its China Post partnership uses its mid-DOF XHand. 陈哲’s comparison: “This year’s ICRA felt a lot like 2025. In 2025, the product everyone remembered was ShaPa’s high-DOF hand; this time it was 五G’s second-generation hand.”
- The second-generation hand has 20 degrees of freedom. For high-DOF hands, roughly 20 is an engineering trade-off that is “basically enough to achieve the various in-palm manipulations humans need.” It fixed 2 major weaknesses of the first generation: thermal management and the inability to backdrive. First-generation joints could not be freely pushed backward, making them vulnerable to impacts and insensitive in force control. The second generation improved backdrivability substantially and thermal management without adding weight. Many visitors tried its dexterity and backdrivability for themselves.
- Size is the hidden deciding factor: 五G’s hand is roughly half the volume of ShaPa’s. Other than 五G and 希诺, most high-DOF hands are 1.5-2x the size of an adult human hand. That means an action a human can learn and execute may not be feasible on an oversized robot hand.
12. 五G’s Niche: Unitree in Dexterous Hands
- 五G’s positioning differs from ShaPa’s. 李帆’s long-term goal with ShaPa is to enter general-purpose embodied intelligence through the hand, with substantial investment in software and algorithms, including a 3-layer model architecture and showpiece demonstrations such as assembling a pinwheel and peeling an apple with 2 hands. 五G is instead “focused on making a low-cost, highly reliable, stable hardware device reliable and durable enough”—like Unitree, becoming infrastructure for frontier research.
- That addresses a core research pain point. After completing VLA research centered on grippers, teams discover that hardware limitations prevent many more complex tasks. Dexterous hands have been developing for 30 or 40 years, yet there is still no particularly mature or inexpensive hardware platform that lets researchers worldwide iterate quickly. Several startups in China and the US released dexterous-manipulation models based on 五G’s hand over the past 1 or 2 months. The core judgment: “Whoever can provide reliable, stable, inexpensive dexterous hands faster will become the industry’s de facto standard faster. More algorithms and software will be designed around your hardware architecture.”
13. The Dexterous-Hand Market: Low-DOF Shipments in the Low Tens of Thousands, High-DOF in the Thousands
- The market has 2 tiers. Low-DOF hands ship in the low tens of thousands per year, roughly in line with humanoid-robot volumes. Players include 零星巧手, potentially the largest by volume; 英石, which uses a linkage design and previously served the disability market; 强脑; 奥翼; and 临界点, spun out of 智源. Their applications are performances and simple pick and place.
- Leading high-DOF hands ship only a few thousand units, represented by 虾爬 and 五G, and target the global dexterous-manipulation research market. Before ShaPa entered mass production last year, there were almost no mass-produced high-DOF dexterous hands available for purchase. Volumes are small, but the technical content gives them influence over algorithms and data, making them a focal point of competition.
14. Genesis’s Dexterous-Manipulation SOTA Is Still 2-Year-Old ALOHA-Style Behavior Cloning
- Genesis was founded in 2024 and started with robot simulation before shifting to dexterous manipulation and full-stack systems more than a year ago. In May it released a dexterous-manipulation model built on a customized 五G hand. 陈哲 sees it as the current SOTA in high-DOF dexterous-hand manipulation. The company says it collected roughly 200,000 hours of data, much of it from heterogeneous sensors, along with teleoperation and real-robot glove data. Demos included solving a Rubik’s Cube, playing piano, and a complete 20-step meal-preparation and cooking task.
- But he places the field clearly: “Today’s dexterous-hand manipulation looks a lot like the real-robot teleoperation we were doing with ALOHA 2 years ago.” With enough high-quality human demonstrations, a system can reproduce actions with high confidence—that is behavior cloning. Change the task or the environment, and success rates can shift materially. Compared with Generalist’s hundreds-of-thousands-of-hours approach to body-free gripper data, “the market still has no particularly clear way to obtain enough high-fidelity data for high-DOF dexterous-hand manipulation.” Pure egocentric video, simulation such as Isaac Sim, gloves, EMG wristbands, and exoskeletons are all in contention, with no consensus yet. What is clear is that a generalized dexterous-manipulation model will require the data problem to be fully solved and unlocked, and the industry is moving rapidly in that direction.
15. Who Owns Dexterous-Hand Data? Unlike Lidar, It Is Hard for Third Parties to Serve Everyone
- The host’s question was where dexterous-hand data will ultimately accrue: third-party data companies, hand manufacturers, or full-stack body makers. 陈哲 first rejected the idea that third parties will dominate. Data is highly dependent on “the hand’s structure and configuration, motor selection, and sensor selection.” Which hand configuration and which number of degrees of freedom are you collecting for? Hands with 22, 21, and 20 degrees of freedom differ in the design of every finger, making transfer and retargeting difficult. The independent data-vendor model that emerged in the UMI era around simple parallel grippers will be difficult to replicate.
- His counterexample is sharp: lidar companies have no real say over the data because point clouds from different lidar systems can be fed through a single pipeline into a larger model. Dexterous-hand data, by contrast, is deeply bound to the hardware architecture. Meanwhile, “Musk has also said that dexterous hands may account for 50% of his company’s engineering effort,” and most companies cannot build SOTA hands, leaving room for third-party hand makers to establish a position. The autonomous-driving analogy is that, aside from full-stack leaders such as Li Auto, XPeng, Li Auto, and Tesla, many companies still buy Nvidia or Horizon Robotics solutions. The conclusion remains open: dexterous hands and humanoid bodies “may be locked in a long-term contest and long-term coexistence.” If broad third-party hand suppliers survive, they will provide much of the data; if not, the data will belong to body makers.
16. A Third Route Beyond Tendon Drive vs. Direct Drive: 希诺’s Hybrid Design and the EMG Side Path
- Another major ICRA release was 希诺未来’s Flex Two hybrid-drive hand. Heavy-load gripping force comes from motors in the forearm via tendons, while miniature motors in the palm directly drive the finger joints. “Our gripping force comes largely from forearm muscles, but small muscles in the palm control fine movements.” The design targets the Achilles’ heel of pure tendon drive: achieving high degrees of freedom requires many tendons to pass through a narrow wrist, making assembly and maintenance extremely difficult. The trade-off is a more integrated commercial model. A forearm-inclusive design “requires deeper coupling and integration with the body maker,” unlike the handless-body approach of Unitree and 智源, which can quickly swap in third-party hands. That helps explain 希诺’s strategic investments from Ideal, JD.com, and multiple other major companies; Chinese manufacturers researching humanoids largely use a Tesla-like tendon-drive design.
- EMG is a side path that has been revived. 原音为 A Region Flow is collecting hand data with EMG wristbands, but “using EMG to control a dexterous hand is not a new idea.” Companies at Waterloo were working on EMG wristbands more than 10 years ago, and 强脑 has also tried EMG control for low-DOF hands. The renewed interest comes from algorithmic progress and much greater capacity to fit large datasets. Researchers want to know whether EMG can reconstruct the position, motion, and force response of every joint in the hand with high precision.
17. Major Manufacturers Follow Tesla’s Tendon Drive for Organizational, Not Technical, Reasons
- 陈哲 has spoken with researchers at many major Chinese companies. Following Tesla is “the straightforward choice”: “Tesla is still pursuing this direction, so asking people to choose a different one creates enormous risk for many engineers. If Tesla turns out to be right, and you try a new direction, who takes responsibility for that mistake?” This follow-the-leader strategy is not limited to dexterous hands; it is the safer choice for large companies across body architectures, algorithms, and many other technical routes.
- Tesla itself is not settled. Musk believes from first principles that tendon drive is closer to human biology, but Optimus Gen 3’s hand “has been continuously iterated”—Gen 3 has been disclosed as 22-DOF and remains tendon-driven, with the timing of mass production unknown. There are also reports that Optimus is testing alternatives, potentially direct drive or hybrid drive. 陈哲’s preference for direct drive “has probably strengthened,” since the latest research breakthroughs have come from all-direct-drive hands. His commercial view is that “an all-direct-drive route without reliance on the forearm may be the only route, or the better route.” A company requiring deep forearm customization will probably become a major manufacturer’s custom supplier rather than an independent standardized-product company. Domestic acquisition exits are also ambiguous: whether you are acquired because the business is failing or because it has become extremely successful is a very different thing.
18. Cosmos 3 Dissected: The First Fully Open-Source Omni Model
- Cosmos 3 can natively process and output text, images, video, sound, and actions across all modalities, “possibly the first time the industry has implemented this in such a way.” Its technical core is MoT, or Mixture of Transformers: an autoregressive transformer handles reasoning, while a diffusion transformer handles generation. The former is better at discrete tokens such as text; the latter at continuous signals such as images and video. For a long time, the 2 models were isolated. Cosmos 3 uses shared attention so their states influence each other, “stitching together 2 very different types of models and unifying understanding and generation.”
- Different tasks activate the components differently. Image-to-text may need only the autoregressive branch, while text-and-image-to-action and video require the diffusion branch. Cosmos previously had no standalone “second-generation” brand: it began as several independent modules—predict, transfer, and reason—released at CES 2025 and later iterated to 2.0, alongside a Cosmos Policy submodel with an action head. This release unifies them in a single architecture. 陈哲 called it “a relatively elegant and ideal architecture in principle” and expects other major companies, including those in China, to develop similar approaches. The larger picture is that Cosmos 3 is optimized for Nvidia’s chip and compute platform, while also promoting Nvidia’s robotics development kit and supporting autonomous driving.
19. A 3-Layer Taxonomy of World Models: From Video Generation to WAM
- Because the term has become overused, 陈哲 used Nvidia’s official June review to narrow the discussion. The bottom layer is the video world model: Google’s Veo and Alibaba’s Wan, built on video-generation models that conditionally generate forecasts of future states; Nvidia places Cosmos 3 in this category. Above that is the action-conditioned world model, which predicts the next state of the world given an action; Dream Dojo, Genie, and Japa belong here. Citing GEAR Lab researcher 高深远, 陈哲 described Dream Dojo as a “world simulator” whose input is your action and whose training data includes real physical changes to the environment. The third category is the world action model, or WAM, which generates both future world states and robot actions and is therefore more of a policy; Dream Zero and Ant Group’s 凌波 VA belong here.
- WAM is the hottest category in embodied AI because it is applied directly to robots. Both host and guest agreed that a text instruction can also be treated as an action on the environment. The difference between a pure video world model and an action-conditioned model is only whether the input is a textual description or a physical action, suggesting that the categories themselves are moving toward convergence.
20. VLA vs. World Models Is a False Dichotomy: To LLM Researchers, “This Is One Thing”
- 陈哲 relayed one of the quarter’s best lines from North American LLM labs: “They do not really understand why we invented VLA and World Model as 2 opposing concepts. They see them as one thing. In their imagination, the best model should be an Omni Model—essentially the Cosmos 3 direction.” VLA is better at generating action instructions, while world models, with video as the backbone, are better at predicting states. “If there is a clever way to combine these 2 capabilities, the ultimate performance will certainly be better.”
- The technical history explains why the idea is taking off now. Video-generation models were already being tried as robot policies in 2023; ByteDance’s GR-1 and GR-2 were early Video Action Model efforts. But there were no good open-source video-generation models, no diffusion policy, and no flow matching. Those matured in 2024 and 2025. VLA could “give a 60-point answer faster,” with Pi 0, open-sourced in late October 2024, the classic example and a route many teams followed. But its essence is behavior cloning, with obvious limits to generalization. Kuaishou’s Kling, Seedance, and Veo 3 now show a high level of physical and common-sense understanding. “In theory, using them for robot prediction and policy should produce very good results. The emergence of this series of elements is why the concept suddenly exploded in 2026.”
21. The World-Model Venture Rush: 49 Companies Counted by April
- The host’s list included Genesis, Liberty AI, Manifold, 逆矩阵, 模式星空, and others, mostly founded in 2025 or 2026. Some became billion-dollar unicorns within the year, with exceptionally young founders. 极佳世界, founded in 2023, currently has the highest valuation. By early April, an investor had already counted 49 world-model companies. Overseas, Ronda AI announced a $450M financing during GDC, while the leader of Nvidia’s Dream Zero team, phonetically “周章,” a Korean, founded a new US company.
- 陈哲’s dissection of the boom is unsentimental: “At bottom, people who made money in big models keep investing, while people who did not make money are buying a ticket.” The underlying drivers are real. Embodied AI is a strategic, long-duration priority with government support, and Chinese companies have strong hardware and model resources. Could this become an electric-vehicle-style rush with hundreds of competitors? “I don’t know, but a major wave of talent and capital is definitely flowing in.” That alone reflects the magnitude of the shift.
22. Pi 0.7: A Lightweight World Model That Lets VLA “Fill in the Blanks”
- 陈哲’s one-line summary of Pi 0.7: it “adds a lightweight world model on top of traditional VLA.” The model predicts images of the outcome over a future interval—a subgoal image—and feeds that prediction back into action generation. He agreed with the host’s intuitive analogy: when throwing a paper ball into a trash can some distance away, you imagine the trajectory. “We adjust our angle and force based on that mental simulation.”
- His overall assessment of Pi is “highly innovative and steadily iterative.” Each version from 0.5 to 0.6 to 0.7 made visible progress. Version 0.6 introduced something akin to long-horizon agent planning, using text to continuously record state and plan actions. “Many small improvements were genuinely industry-leading at that point.” Many Chinese teams have fine-tuned Pi after its open-source release or used it as a benchmark, underscoring both its influence in setting industry standards and the scarcity of independent pathfinders.
23. Gen One: 500,000 Hours of Self-Collected Data, 99% Success, and a Separate Route
- Gen One has generated the most discussion among Chinese practitioners. 陈哲 cited 3 breakthroughs: execution speed is materially higher and no longer “slow and leisurely”; the company “claims” high accuracy on complex long-horizon tasks, with several tasks rising from an average success rate in the low 60% range to 99%; and, most importantly, Generalist collected 500,000 hours of UMI-style, body-free real-interaction data and trained an end-to-end model from scratch without a pretrained base. Its report clearly shows a sharp improvement in generalization and capability from 10,000 to tens of thousands to 500,000 hours of data. “By Generalist’s own claim, they have found the key to scaling.”
- Why is training without a pretrained base so unusual? “The research community is prone to path dependence. Independently creating a new technical approach and collecting and training on that much data requires considerable courage and determination.” That connects to his observation about US venture culture: “It is hard to find 2 or 3 teams telling exactly the same story.” Without clear differentiation and innovation, it is difficult to gain momentum, market recognition, or an acquisition—an acquisition is a normal exit for US hard-tech startups. On labels, Peter Florence wrote explicitly that the approach “is neither pasting actions onto a VOM to turn it into a VOA nor a world model.” It trains a transformer entirely on native physical-interaction data, with roots in the group’s earlier RT-1 work at Google, before RT-2 set off the VLA wave.
24. Google Wants to Build Android for Robotics, and the Endgame Question Applies to Every General-Purpose Model Company
- Gemini Robotics’ ER 1.6, or Embodied Reasoning, released in April, did not make the top 5 because it is not a policy model. It is a general-purpose model that improves spatial and task understanding and reasoning, and its report directly benchmarks it against a general model such as Gemini Flash 3.0. The partnership ecosystem reveals the intent: Boston Dynamics’ Spot, with several thousand units shipped for oil-and-gas inspections, stair climbing, dashboard reading, and anomaly checks; Apptronik; and 思灵’s humanoid AgiOne. “They want to be the Android of this field,” occupying a brain- and API-oriented position.
- 陈哲 pushed the question to first principles: “What reason do we really have to believe that the embodied-AI model, many years from now, will not be provided by general-purpose model companies such as Anthropic, OpenAI, or Google?” If the Omni Model route works, could all of a robot’s spatial understanding and action prediction be fused into a larger general-purpose model? The standard counterargument is that general model companies do not build hardware and cannot access the data. His response is capital: big-model companies spent tens or hundreds of billions of dollars on corpus construction and data labeling, mostly for text and images. Their weakness today is simply that they have not optimized for this task.
25. OpenAI Restarts Robotics: The Sora Team Turns Around, and 10K GPUs Are Only a Rounding Error
- OpenAI officially announced its robotics team at the end of May, recruiting top full-stack talent across hardware, systems, and machine learning. The short-term goal is to build AI-compute infrastructure for itself; the long-term goal is robots anyone can use. Team leader Arditia Ramesh is the author of DALL·E and joined OpenAI straight out of university. The team’s Twitter name is “World Simulation Research.” 陈哲 confirmed that OpenAI’s robotics team was built from the group that worked on DALL·E and Sora, and that “OpenAI has always viewed Sora internally as a world model.”
- How large will the investment be? 陈哲 offered a number that resets industry expectations: “Cosmos 3 probably already represents the highest compute budget for embodied-AI training today, roughly at the 10K-GPU level. That is already the ceiling for compute in embodied research.” Ant Group’s 凌波 team is also around the 10K-GPU level. By that standard, this is “not an especially large compute budget” for leading LLM companies such as OpenAI and Anthropic. One thing is clear: robotics research remains early and requires hardware validation. “What I know is that OpenAI will invest particularly heavily in dexterous hands.” Once it enters the market, it cannot keep working only on grippers. 陈哲 expects OpenAI to deliver SOTA progress in the field over the coming period.
26. Unitree Anchors Valuations, While the Model Endgame Polarizes: “Anthropic Became Anthropic”
- Unitree’s STAR Market IPO passed its hearing in Q2, “setting a valuation anchor for all leading embodied-AI companies today.” The body market may not be winner-take-all: physical limits around power density, batteries, weight, and volume will produce different humanoid forms, from smaller units for entertainment, performances, and research to larger units for heavy loads, low speeds, and payloads. The brain side, however, “will at least be oligopolistic.” “Intelligence will inevitably polarize: either pursue the best intelligence or settle for acceptable intelligence.” With low marginal model costs and open-source sharing, intelligent models will likely divide over the long term into “high-quality open-source models and top-tier closed-source models.” If a closed-source model cannot surpass the SOTA open-source model, “it has no reason to exist.” The likely open-source suppliers are ecosystem companies such as Nvidia that monetize compute rather than models.
- Model startups therefore face 2 ultimate questions: how to compete with the best open-source models when the suppliers do not depend on model revenue, and how to compete with the oligopoly of closed-source models from Google, Alibaba, ByteDance, and Anthropic. Has any founder offered a convincing answer? “Investors repeatedly asked Anthropic the same question when it was founded, but Anthropic became Anthropic.” Being unable to answer at the start is not decisive. But there is a reality check: when OpenAI was founded, it did not face a crowd of language-model companies the way embodied-model startups do today. This embodied-AI boom was itself propelled by the IPO and profit windfall of the previous language-model wave. “If everyone thinks this way, the market will inevitably see excessive competition and excessive investment. The ultimate winners will still be relatively scarce.”
27. China-US Differences in Commercialization Pace: “The Posture and the Facts Are Still Quite Different”
- The pattern at US frontier companies is a group of visionary scientists sustaining investment in a field with huge potential and long-term uncertainty. OpenAI had no obvious commercialization for 7 years from 2015 until ChatGPT, which did not stop it from attracting talent and resources. Physical Intelligence is following the same pattern, repeatedly releasing and even open-sourcing leading models, but “so far we still have not seen a particularly clear stage of commercialization.” When the host asked whether Figure was not supposed to be deploying, 陈哲 responded that Figure still had “no specific, genuinely validated deployment revenue. The posture and the facts are still quite different.” Asked whether China’s posture is fact or theater, he was direct: China has very little tolerance for exploring highly uncertain, very long-term opportunities, so companies must project that posture—and “preferably make it a fact.”
- How can a company secure resources in China while saying it will not commercialize for 10 years? The host had a neat answer: “You are 梁文锋.” 陈哲 agreed that DeepSeek is rare: “When you are successful enough, you have enough conviction in the thing.” OpenAI and DeepMind were also originally the projects and investments of billionaires. The bigger disagreement is the cycle. If one could see clearly that broad commercialization would take 10 or 15 years, as with autonomous driving—“Waymo has achieved commercialization, but it is not yet profitable”—would anyone still invest? His view: “Compared with last year, this cycle is definitely shortening.” Compute, data, and capabilities from language models are moving into embodied AI, accelerating the field in fact. Figure’s and 星动纪元’s parcel-sorting deployments are evidence that practical deployment is possible.
28. Outlook: The Next 6 Months Are the Last Window for Major Manufacturers, with Full-Stack and Vertical Models in Long-Term Competition
- The core call for next quarter is that humanoid robotics will become more mainstream as a form factor and platform, more products will land in specific use cases, and more major companies will announce humanoid programs. From announcement to product takes at least 1 to 1.5 years. The more aggressive timeline comes from the embodied-AI founders 陈哲 speaks with; his own formulation is: “The next 1 or 2 quarters, or the next 6 months, are the last window to enter the game. If you do not start in 2026, by the time humanoid robots begin broad deployment in 2027 or 2028, you will have no position.” Internet giants such as ByteDance, Tencent, and Alibaba are more likely to follow the Pi and Anthropic model and build intelligence-focused foundations, since “the business model of the model is highly connected to their own capabilities.”
- Chinese embodied-AI startups face a double paradox: they struggle to build companies with no revenue for 10 years, while investors are even less willing to tolerate a company that is not full-stack. The result is that “every embodied-brain company is announcing that it is building a robot body.” 陈哲 cited the unconventional view of 许华哲—that some work should be left to the ecosystem rather than brought entirely in-house—as “possibly correct, and a very American idea,” but immediately added the condition: if you do not go full-stack, you must be the best in your niche to have value. The endgame remains unresolved: “There will probably still be opportunities for full-stack companies to become the ultimate winners—companies such as Apple, Xiaomi, and Huawei will exist in the embodied era. But interestingly, the largest companies by market cap today are not full-stack: Nvidia is not, and if Anthropic or OpenAI eventually goes public, they may not be considered full-stack either.” Full-stack versus vertical integration will remain a long-term contest. Even Unitree faces investor criticism that it is “just a hardware company with no AI.” 陈哲 rejects that framing: “Even if Unitree remains a hardware company, it can still be a very profitable hardware company. I see no reason it cannot grow into software or AI if it genuinely wants to do so strategically.”
Verification Notes
- The RAW contains different phonetic renderings for 星动纪元/心动纪元 and 五G/沙帕; this edition retains the forms directly supported by the RAW.