SOEs Buy In at Peak WRC: Bubble Squeezed, Embodied AI Goes Real-World
SOEs Buy In at Peak WRC: Bubble Squeezed, Embodied AI Goes Real-World
Summary
- The host framed this year’s WRC as the industry’s shift from stage demos to real-world deployment. 叶杨笙 saw more visitors than last year—“you could barely move; it had caught up with WAIC”—along with more machines actually running. Products with little progress were no longer easy to bring out, let alone put on display. Timed demos may signal a lack of confidence in success rates or hardware that cannot support extended operation; 仙工咖啡, 自变量’s logistics-sorting system and several 星尘智能 demos ran continuously.
- Shipments are staggering, but the share backed by real demand may be very small. The program cited estimated global humanoid-robot shipments of roughly 22K units in 1H 2026, with Chinese manufacturers accounting for 97%; 智元 shipped 9,700, 宇树 more than 7,000 and 银河通用 1,100, while the top 5 accounted for 86%. 叶杨笙 believes that, excluding dancing and other performances, a 10% share of genuine deployment would already be an excellent result. Space-capsule retail does not qualify because of weak ROI, poor generalization and limited data value; other machines are used for data collection, research or education, while many may simply be gathering dust in factories.
- There is substantial hype in the “VLA + world model + full-stack in-house development” narrative. 叶杨笙 says there may be only 1 or 2 startups currently buying more than 1,000 GPUs, aggressively purchasing data and investing in pretraining; he later stressed that this was incomplete-counting territory and that the true number with thousand-GPU commitments may be just 2 or 3—far below the PR narrative. A thousand-GPU cluster plus storage can cost several hundred million yuan a year, a figure that can be reverse-engineered from the financials and roadshow materials of companies lining up to list, but those materials do not show anything close to that many firms doing it. Some leading companies are keeping cash on the balance sheet until the paradigm becomes clearer; 叶杨笙 sees nothing wrong with that from a pragmatic standpoint.
- The gap between speed and ROI remains substantial. On claims that robots now reach 70%–80% of human speed, 叶杨笙 says logistics sorting may be close, but most tasks are not; even 50% could produce acceptable ROI if the machine can work continuously. Folding clothes, however, still takes a robot 2 minutes versus 10 seconds for a person—a gap that remains “an order of magnitude.” The industry is improving, but not at breakneck speed.
- Brain-side capabilities are highly commoditized, making Agent OS one of the more pragmatic paths forward. Most companies are following π0 and iterating on its results; a booth demo alone cannot reveal whether a system uses a world model or proprietary pretraining—you have to take the vendor’s word for it. 叶杨笙 favors atomizing classical control and model capabilities, then letting an Agent schedule them for long-horizon tasks, recovery and disturbance rejection; latent-space prediction by world models is inherently difficult to show.
- The underlying bottlenecks are the “nervous system” and data standards. By “nervous system,” 叶杨笙 mainly means a high-speed, reliable bus and protocol stack connecting sensors and actuators; robotics still lacks a common standard. Data problems span formats, timing, labeling and collection conventions. Multiple sensors need synchronization at the tens-of-nanoseconds level, ideally through hardware timing. 仙工 plans to build datasets from scene reconstruction and assetization, human demonstrations, real-robot trajectories and human takeover data, alongside benchmarks, real-machine feedback and training/evaluation infrastructure.
- This year’s gains came mainly from post-training; next year’s test will be data and VLA. 叶杨笙 says the progress did not come from world models but from post-training, better data quality and stronger engineering, with AI Coding adding further leverage. Data and VLA may produce more meaningful results next year; the acceptance tests are movement speed and whether a robot can move through space and work continuously from start to finish without relying on timed demos. 宇树’s listing, the financing cycle and the long duration of the robotics market all mean companies still need enough capital to survive the cycle.
Deep dive
1. WRC Gets Hotter, but Products with Little Progress No Longer Belong on the Show Floor
- The host sees this year’s WRC as a point at which the industry is moving from performance to practice. 叶杨笙’s read from the floor was that attendance was even heavier than last year—“you could barely move; it had caught up with WAIC”—and that more machines were genuinely operating. Products with little progress may have had “no way to show up, and no face to put them on display”; naturally, few visitors came to see works still in progress.
- His read on timed demos is twofold: the team may “not be confident enough in the success rate” and need someone standing by, or the hardware may not be able to support long-duration operation. 仙工’s semi-automated coffee machine, by contrast, runs the full workflow from tamping and extraction to knocking out the puck, with a process close to how a person works—“making coffee all day long.”
- 自变量’s logistics-sorting system ran all day at a fast takt time, while several 星尘智能 demos also appeared capable of running continuously. Asked whether the show left him more confident or more nervous, 叶杨笙 said: “Relatively speaking, I feel pretty confident.”
2. 星尘智能’s Supermarket Picking: One of the Hardest Demos on the Floor
- The demo that stayed with 叶杨笙 most was 星尘智能’s supermarket picking task. The robot had to put products into a shopping bag, pull out the drawstring, fasten it with a machine, and seal the bag. The precision requirement was high and the system could fail multiple times along the way, but it ultimately completed the task. In his view, it was the longest, hardest and most technically demanding task in the venue.
- 自变量’s logistics sorting ran at a rapid takt time. Generalized grasping has already been achieved by several companies and can be used to show the VLA reasoning process; it is already usable in a meaningful range of settings.
- Robots running and sprinting, along with the 天工机器人运动会, were among the show’s hot topics. 叶杨笙’s reaction was that they were “perhaps running so fast they were throwing sparks.” 加速进化’s soccer interaction lets visitors play goalkeeper, but after trying it himself he found it difficult to defend against, and the robot’s shots carried plenty of force. The table-tennis robot was mainly a showcase for low-level motor control; completing the full task remains very difficult.
- On the discussion of 黄仁勋’s daughter visiting WRC, 叶杨笙 said he did not see her because he was “too busy.”
- In dexterous hands, the host highlighted 章鱼动力, founded this year. 叶杨笙 noted its integrated hand-and-forearm design, which looked highly biomimetic. Electromyography is another emerging direction: an EMG band worn on the wrist, combined with a sensing glove, can provide two independent data streams to validate hand posture, while also reading muscle activation in reverse to verify that motion capture is working correctly.
3. An Insider’s Field Guide: Teleoperation or Autonomy?
- When walking the floor, 叶杨笙 watches the smallest movements and looks for hidden teleoperation consoles. If one arm of a dual-arm robot moves and then the other moves, someone may be teleoperating it. If both arms move at once along different trajectories, autonomous planning is more likely.
- Recovery after failure is another key test: does a human restart the system and begin again, or can it remove the object that caused the error and continue? That tests both model capability and the ability of long-sequence attention to handle extended operations.
- He will interact with a pure-white object whose color is close to the tabletop, then use his phone’s flash to alter the lighting and observe how the system responds to changes in the object and environment.
- The installation period is also revealing. After the environment or lighting changes, some companies’ success rates initially fall sharply; they have to collect more data on site and retrain. 叶杨笙 saw many teams recording data beside their booths for 2 or 3 consecutive days. The host had previously cited 智元 as an example of a company with people constantly demonstrating and interacting with visitors; 叶杨笙 then noted that 智元 is fully simulation-based and exhausts possible environmental conditions in simulation, so it does not need to collect supplemental data on site.
- Other useful questions include: “How many clips did you cut, and how many times did you speed them up? Can it still move if the network goes down?” And: “How many attempts and scenarios went into the success rate?” The latter is fundamentally a question of how the denominator was chosen. 叶杨笙 believes that selective disclosure of data definitions is understandable in a hot, fiercely competitive market.
4. Procurement Day Pushes Embodied Intelligence into Real-World Use Cases
- This year’s conference added a Procurement Day, bringing central SOEs and large companies to the floor with purchasing projects in hand; customers for embodied intelligence were present as well. The host sees the back half of “human-machine coexistence, industry-academia co-prosperity” as a major difference between this WRC and previous editions.
- 仙工 demonstrated a hotel food-delivery scenario combining a delivery robot with a robotic arm. The arm picked up the order from the floor, pressed the elevator button, and finally placed the food from the robot onto the floor. Traditional delivery robots generally require a courier or front-desk worker to load the food, then use elevator controls to reach the vicinity of the room.
- The host noted that the previous generation of hotel robots already integrated with elevator controls and asked why they needed to press the button themselves. 叶杨笙’s answer was cost and compliance. Retrofitting elevators requires additional investment; overseas, safety regulations and even higher retrofit costs effectively rule out the approach. Robots therefore need to use elevators as much like people do as possible.
- Elevators in China are classified as special equipment, so modifications require filing and construction work, adding both time and cost. The arm-based approach also reduces the need for staff to load and retrieve orders. Acceptance may be higher in hotel apartments, residential compounds and other settings with limited staffing.
- Last year, humanoids were commonly shown standing still while their lower bodies remained fixed. This year, more companies put 2 dual-arm systems directly on tables for demos. 叶杨笙 sees no need for a humanoid to stand and make coffee if its lower body does not need to move. The host argued that the humanoid form could be a milestone toward the end state of machine intelligence, but that being humanoid does not necessarily require 2 arms; some tasks can be handled with 1.
5. The Shipment Myth and the Reality of Demand
- The program cited estimated global humanoid-robot shipments of roughly 22K units in 1H 2026, with Chinese manufacturers accounting for 97% of the global total. 智元 shipped 9,700, 宇树 more than 7,000 and 银河通用 1,100; the top 5 accounted for 86%. 叶杨笙 believes China’s embodied-intelligence industry already holds a dominant position globally, thanks to the strength of its manufacturing base and supply chain. He also said the model capabilities demonstrated by overseas companies such as 派, journalist and Diana are genuinely ahead, but their humanoid hardware is difficult to match against Chinese products.
- On genuine demand, 叶杨笙 repeated last year’s judgment: if dancing and performances are excluded, a 10% share of actual deployments would already be “very, very good.” In his view, reported order volumes should be “divided by 10.”
- He does not regard space-capsule retail as genuine demand. First, the ROI does not work. Second, the task has weak generalization: the environment is fixed, the products are fixed, and the setting cannot keep generating valuable data. Nor can a company deploy more robots to spread costs and improve the generalization of the underlying intelligence.
- Beyond real customer demand, some robots are used in data factories, research and education, or R&D testing; many “may indeed just be gathering dust in factories.” Shipment figures also depend on revenue-recognition rules: is revenue booked when the unit ships, or only after the customer signs acceptance? Private companies do not disclose complete financial statements, making it difficult for outsiders to assess the numbers accurately.
- 叶杨笙 added that 仙工 already has enough factory customers and use cases, but he has not yet seen humanoid robots actually deployed at those customers.
6. Speed and ROI: The Gap Is Still an Order of Magnitude
- On 星尘智能’s 高继祥 saying robots operate at roughly 70%–80% of human speed, 叶杨笙 said the answer depends on the task. Most tasks are not there yet; logistics sorting may be close.
- He believes 70%–80% is already enough to make the ROI work because machines can operate continuously while people need breaks. Even 50% of human speed could be acceptable in some settings.
- On 熊友军 of the Beijing Humanoid Robot Innovation Center saying the ROI gap for humanoids is narrowing rapidly, 叶杨笙 said he had “not seen any obvious narrowing.” Folding clothes is one example: a robot may need 2 minutes, while a person can finish in at least 10 seconds. The gap remains an order of magnitude.
- The host clarified that 熊友军 may have been talking about the future trend, and 叶杨笙 agreed. From WIC to WRC, he has clearly felt the technology improve versus last year, but not enough to conclude that the industry is advancing at “breakneck speed.”
7. Use Cases Converge on a Red Ocean: Transport and Machine Tending
- Logistics transport, industrial loading and unloading, coffee machines, object storage, teleoperation replication and generalized grasping have become the most common demos. 叶杨笙’s explanation was blunt: “Perhaps these are the only scenarios the technology can handle; it has not reached the level required for anything more complex.”
- As factory automation has matured, most jobs no longer require people. The tasks that still do are mainly moving things and loading or unloading machines. Screwdriving has largely been taken over by collaborative arms or dedicated equipment and is no longer a primary entry point for embodied intelligence.
- Industrial manufacturing, automotive and 3C electronics are highly digitized and automated, but they also have sufficient budgets. 叶杨笙’s past industrial customers included logistics and semiconductors. These markets may be red oceans, but their scale is close to 1 million units, and the demand itself is real.
- More than 300 companies in China are competing in similar directions, which 叶杨笙 sees as normal: the opportunity is limited and the participant pool is too large. Based on the evolution of mobile robots and collaborative arms, the end state may consist of leading companies plus some backed by central ministries, with perhaps only a dozen or so survivors.
8. World Models Remain Difficult to Show
- 叶杨笙 believes there are few world-model demos because world models are intrinsically difficult to present. He cited 李飞飞’s three-part framework: generation, simulation and evaluation, and planning.
- The first 2 categories are more generative: natural language is used to create physically coherent, interactive scenes for simulation or training. A planning-oriented world model may exist in latent space, predicting how objects or the environment will change over a future interval and using that forecast to support robot planning.
- To show a latent-space prediction, teams generally have to diffuse it back into a visualizable space. The resulting video may look little different from a VLA demo.
- “VLA + world model + full-stack in-house development” is now standard language across technology companies. 叶杨笙 sees that as normal during the exploration phase. In key modules, embedded software and drivers must be developed in-house to control the robot and integrate with low-level motor control; leading companies can also pretrain their own VLA and world models.
9. Pretraining Investment Remains Cautious
- On whether there is hype in “VLA + world model + full-stack in-house development,” 叶杨笙 said the “water content is extremely high.” He defines genuine pretraining as buying more than 1,000 GPUs, continuously purchasing large volumes of data and actually putting the resources into pretraining.
- Excluding ByteDance, Ant and CATL, he estimates that only 1 or 2 startups may be doing genuine pretraining. He later emphasized that this was an incomplete count and that the true number with thousand-GPU commitments may be just 2 or 3—far below the PR narrative.
- The annual cost of a thousand-GPU cluster and storage can be reverse-engineered and may reach several hundred million yuan on an equivalent-H100 basis. More than 20 companies waiting to list already have to submit financial and roadshow materials, but those materials do not show anything close to that many companies running thousand-GPU pretraining. China also does not have enough GPUs to allocate to that many companies.
- Many companies are not short of cash; they simply do not yet understand the paradigm and worry that current spending would amount to “making the wedding dress for someone else.” Once employees move, a competitor could reuse the path already explored. As in autonomous driving, once technical routes converge, competition may shift toward data feedback loops and engineering capability.
- The discussion turned to the idea of the “perpetual company”: after raising several billion yuan, a company makes little investment for the time being, and annual interest income may cover the salaries of hundreds of R&D staff, allowing it to wait for an inflection point. 叶杨笙 described the strategy as “the mantis stalks the cicada, unaware of the oriole behind,” and said it was not wrong from a pragmatic standpoint.
- The host cited 王兴兴’s view that if a company does not build its own feel for data, architectures and model algorithms early on, simply waiting for the paradigm to mature and then copying it may not work very well. 叶杨笙 agreed that the point had merit, but said robotics also requires real-machine data and scene understanding. The tension is that without pretraining, a company may not know what good data looks like; yet data collected in advance may lose value when the paradigm changes. When the inflection point arrives, it also matters whether the company already has enough machines deployed in the field. That is why companies are competing for a limited pool of scenarios and customers.
10. Agent OS: A More Pragmatic Brain Architecture for Now
- 叶杨笙 believes brain-side capabilities are currently highly homogeneous, with most companies following the global frontier. The industry broadly recognizes π0’s performance, and many companies are building their own improvements on top of it.
- A booth demo cannot reveal whether a system uses VLA, VLA plus a world model, or proprietary pretraining: “You can’t tell; you can only listen to what they say.” Booth tasks are usually limited and cover relatively small areas, where VLA alone may be sufficient.
- 叶杨笙 mentioned a recently released Google architecture based on an Agent System, with multiple expert models and engines handling orchestration, though he did not recall the model’s name. Rather than continuing to train one enormous model, capabilities can be atomized: some functions can be implemented with classical control algorithms for greater stability and efficiency, while others can be handled by models.
- Agent OS would understand what capabilities are available, schedule them to complete long-horizon and complex tasks, and recover from errors or withstand disturbances. 叶杨笙 sees this as a pragmatic approach for now; 仙工 is conducting related research.
11. The “Nervous System” and Data Standards Are Foundational Bottlenecks
- By the “nervous system,” 叶杨笙 mainly means the high-speed, reliable bus and protocol stack connecting sensors and actuators. Once a robot has a large number of both, it must receive, process and transmit complex signals quickly while maintaining interference resistance and stability.
- The analogy is autonomous driving: as the number of sensors such as lidar and cameras rises, a vehicle’s existing CAN bus or EtherCAT may fall short on performance and bandwidth, with Ethernet-like solutions potentially needed in the future. Robotics likewise lacks a unified bus or protocol standard for connecting sensors to the low-level motor-control layer and the high-level brain. The problem spans both the chip layer and protocol-software layer.
- The 4 major data-standard pain points are: (1) format—each company defines its own data structures, so data cannot be used directly by another company’s model and must first be converted, an expensive piece of “dirty work” at scale; (2) timing—data from multiple sensors and actuators must be timestamped, with synchronization generally needing to reach the tens-of-nanoseconds level, ideally through hardware timing because software-only synchronization lacks sufficient precision and can create timing errors that degrade model performance; (3) labeling—there is no industry consensus on whether “this is a cup” is enough or whether transparency, physical state, speed and other attributes must also be labeled, making labels difficult to reuse; and (4) collection conventions—the industry is still discovering what data the brain actually needs, while differences in each company’s embodiment make their data formats, labeling systems and timing conventions difficult to normalize.
12. The Lossiness Problem and 仙工’s Data Blueprint
- The host relayed 王兴兴’s view from the opening speech, which 叶杨笙 called “very right, and very precise”: a language model’s inputs and outputs are numerical encodings in a fixed vector space, with relatively little loss; every input and output in a robot introduces deviation and loss, and subtle errors at the tactile level may be difficult to correct, driving down the success rate.
- 叶杨笙 attributes the problem to incomplete data. It is difficult even to specify how many modalities humans rely on to complete a task, while robots currently collect too few, relying mainly on images and vision and only recently adding tactile or force information.
- Tactile sensing, EMG signals and electronic skin could become incremental directions. Papers on VLA systems with tactile or EMG inputs are increasing, but 叶杨笙 has not yet seen a genuinely deployed case. Data-collection gloves and equipment remain expensive, making real-world use difficult.
- 仙工 hopes to build a relatively complete dataset. First, it will capture environmental scenes, perform 3D or 4D Gaussian reconstruction, and turn foreground objects such as cups, cabinets, doors and drawers into interactive physical assets; reconstruction alone produces pixels, not objects that can actually be grasped or manipulated. Second, it will collect wide-angle human demonstration data using heterogeneous or specialized equipment, with long-horizon tasks and semantic labels. Third, it will collect real-robot trajectories from teleoperation, existing models, classical methods or planning and control. Fourth, it will collect human-takeover data in which people intervene, demonstrate or teleoperate after a robot makes a mistake, showing the robot the correct action; 叶杨笙 believes this may be the highest-quality data of all.
- These data can train VLA and world models and support benchmarks that define tasks, success rates and success-state evaluation criteria. 仙工 plans to open the dataset, let models compete on a leaderboard, and connect real robots to the platform for live data feedback. The back end would handle model training, simulation and evaluation before returning the model to the robot’s brain.
- 仙工’s data business aims to reuse the underlying capabilities it once used to “make controllers number one globally,” extending both ends of the stack: connecting more real machines and continuously feeding back data at the front end, while handling training, simulation and evaluation at the back end. The goal is to help the industry establish a de facto standard through a value-generating system.
13. Post-Training, Capital Cycles and Next Year’s Metrics
- Asked whether this year’s technical progress came from reinforcement learning or world models, 叶杨笙 answered: “Not world models; it came from post-training.” He also pointed to better data quality and stronger engineering, and agreed with the host that AI Coding would accelerate engineering capability.
- 叶杨笙 believes the industry is still exploring everything from data pipelines to embodiment-free data and world models, with no especially large results yet. But companies devoted substantial resources this year to finding high-quality, low-cost data and building pipelines and infrastructure, so next year “should produce many surprisingly good results” in data and VLA.
- When he watches WRC next year, he will focus on 2 indicators: whether action speed is fast enough, and whether the robot can move to different positions in space and continuously complete different tasks, operating from start to finish rather than relying on a timed demo.
- On 宇树’s listing, 叶杨笙 said it would be good for robotics: the higher its market cap, the more it can establish an industry benchmark and build confidence. He did not directly judge whether the revenue of companies waiting to list involved financial problems, but explained what an audit might examine: a robot placed in a factory but not used cannot support revenue recognition; a signed acceptance form is also insufficient if the robot is not actually used; a sale to a distributor does not count if the unit has not reached the end customer; and if a company sells robots to a data factory and then buys the data back, that may constitute “internal circular trading” and pose a problem in the strict sense.
- On financing, the host mentioned the funding-starvation effect in which an investor who comes to a company is left waiting for several days, or companies raise a new round every month. 叶杨笙 did not criticize the behavior and instead said: “Raise as much money as you can, quickly.” Robotics and embodied intelligence are long-duration markets, while capital markets are cyclical; companies may need enough runway to survive periods when funding is difficult to obtain.
- 叶杨笙 has been involved with robotics since around 2013, starting with RoboCup soccer robots, then moving into mobile robots and controllers before entering embodied intelligence. Costs have fallen sharply, while performance and intelligence have improved substantially: a small motor that once cost RMB10K can now fit into a humanoid costing around RMB100K, or even less than RMB100K. The gap between companies may be no more than 6 months, but the industry’s overall starting point remains low and the room for improvement is still large.
- The image he remembers most vividly was 仙工’s lottery draw: the crowd was enormous, and once the barriers were opened, people unexpectedly formed a queue along them.