Talk with 王鹤: Embodied AI’s Fringe History and Capital-Fueled Disorder
Summary
The test for embodied AI over the next 5 years is not the demo, but whether leading companies can achieve autonomous productivity at the scale of “10,000 units in a single year.” 王鹤 believes that if Chinese and US peers still cannot clear 10,000 units by then, the industry cycle will lengthen sharply and capital confidence may turn cold: “Our field has been falsified; the bubble was all bubble.” That is the hardest milestone against which today’s lofty valuations can be judged.
银河通用 is betting on “productivity-grade products”: start with proven wheeled bases, dual 7-DOF arms and harmonic reducers to make mobile manipulation—moving, picking and placing—a replicable solution, then expand the skill set step by step. Founded in May 2023, the company’s strategic round is closing and its valuation has surpassed $1B. It plans to mass-produce at the 1,000-unit level in 2025; 王鹤 believes the current hardware, once converged, can support at least 100,000 units, while robots in unmanned pharmacies are already running at more than 20 hours a day.
Real-world data still cannot independently support a general-purpose VLA: the manufacturing cost alone for 10,000 full-size humanoid robots would be at least RMB1B, while multi-shift teleoperation, labeling and quality control could push monthly maintenance costs into the hundreds of millions or even RMB1B. Real-world data accounts for roughly 1% or less of 银河’s training set, with most of the work handled by its in-house synthetic-data pipeline. The logic is not to reject real data, but to use synthetic data to reach roughly 98% success first, then repair the remaining errors through deployment feedback and reinforcement learning.
VLMs are weaker than LLMs, and VLAs are harder still; 王鹤’s core explanation is data coverage, not the claim that “language is the essence of intelligence.” Internet text covers a large share of what humans might express, while public images capture only a small fraction of what human eyes see, and robot action data has only begun accumulating over the past 2 years. A fully general VLA within 1 or 2 years is therefore impossible in his view; the realistic path is sufficient generalization within specific industries first.
The commercial dividing line is not the number of skills demonstrated, but whether marginal development costs fall materially when the robot is deployed at the next customer. At one extreme are years of “long-term floating,” with orchestration or teleoperation used to film videos but no routine output; at the other are project-based labor shops that redevelop everything for every material, factory or store. Even several hundred million yuan in revenue is unlikely to produce high-quality scale. 银河 is building one capability that can replicate across brands, stores and new products in shelves, pharmacies and industrial sorting.
The industry is burning through its credibility; the basic verification sequence should be “demonstrate publicly, with teleoperation prohibited,” followed by daily real workload and long-term operating records. 王鹤 uses Figure’s high valuation to frame both sides: if its robots can truly enter production lines reliably, long-term value could reach several trillion dollars; if there are only a dozen or 20 units, no routine operation and a workflow different from what the company claims, a valuation of tens of billions of dollars lacks productivity evidence.
China’s push for embodied productivity is not just an industrial opportunity but a time window created by labor constraints. 王鹤 worries that aging and declining birth rates could reduce available labor over the next 10 to 20 years to less than half of today’s level. Even combining the populations of China, Japan and neighboring countries would not fill China’s future labor gap. Home robots will first appear in small batches as products that do “light work, with heavier emphasis on companionship and entertainment”; he expects pilots within 3 to 5 years, while true generality remains a decades-long evolution.
Deep dive
1. Embodied AI in China Went from “Unsearchable” to a $1B Valuation
银河通用 was founded in May 2023. 王鹤 says its strategic round is closing and its valuation has surpassed $1B; he was 33 at the time. 张小珺 judges that his MBTI is probably ENTJ. 王鹤 is simultaneously the company’s founder, CTO and an assistant professor at Peking University.
王鹤 received an offer from Peking University at the end of 2020. After returning to China in 2021, he named his lab the Embodied Perception and Interaction Lab, calling it one of the earliest mainland labs to use “embodied” in its name. At the time, “embodied AI” was barely searchable on the Chinese internet.
Around 2022, the Beijing Academy of Artificial Intelligence asked him to lead a study on whether the field was worth investing in. The discussion concluded that existing internet data could train capabilities in virtual digital environments, but intelligence in the physical world required data generated through interaction between a body and its environment. BAAI subsequently established an embodied AI research center.
2. Early Embodied AI Was Defined by the Source of Knowledge, Not Robot Type
王鹤’s academic retrospective uses the distinction “Internet AI versus Embodied AI”: the former mines knowledge from existing internet data, while the latter generates data through interaction between a body and its environment, either in a physics-compliant simulator or in the real world.
“Embodied AI” became a consensus term only relatively recently. He recalls that computer-vision workshops around 2018 or 2019 still more often discussed embodied agents. It was not until 2020 that US academia began using the contrast between Internet AI and Embodied AI consistently.
In 2020, he organized the Simulation Technology for Embodied AI workshop with 苏昊, 易立 and other researchers to discuss how synthetic simulation could advance embodied intelligence. By 2021, the term was already widely used in the US, while it was only beginning to appear in China.
3. Embodied AI Began as Computer Vision’s Revolt Against Passive Perception
The concept was driven mainly by computer-vision researchers. ImageNet, recognition and semantic segmentation relied on internet images and human annotations. Once that paradigm began to settle, the natural next step was to “empower robots with vision”—moving from digital vision to robotic vision.
Traditional robotics was fragmented across Mechanical Engineering, control, CS, EE and even aerospace, with different groups working on hardware, control, path planning, filtering and embedded systems. 王鹤 cites an old saying: “Robotics has no independent scientific problem of its own.” It is more like putting multiple disciplines into one system.
Early work using reinforcement learning to control quadrupeds therefore did not necessarily call itself embodied AI. To control researchers, neural networks, MPC and PID were simply control methods, with inputs still consisting of joint angles, IMU readings and other proprioceptive information. The new narrative emerged when vision moved from passive perception toward active observation and interaction with the environment.
4. Language Elevates Intelligence but Is Not a Necessary Condition for It
张小珺 asked whether “language is intelligence” and whether vision itself produces no intelligence. 王鹤 disagrees. Insects evade attacks; animals forage, avoid predators, navigate and control their bodies. Today’s powerful VLA models may not match those abilities, but the inability to speak does not make an organism unintelligent.
His definition is straightforward: intelligence is “responding appropriately to circumstances.” When the environment presents different challenges and needs, an intelligent system acts to achieve its goal. An insect dodging a hand is a low-dimensional, short-horizon reaction; a person decomposing a career crisis is a high-dimensional, long-horizon reaction. Both are fundamentally interactions with the environment.
Language is more like a major jump in advanced intelligence. It compresses communication into a one-dimensional sequence, helping transmit knowledge, organize thought and plan over long horizons, but it is not the essence of intelligence. Vision is also just one sensor; some organisms do not use eyes and still sense their surroundings through finer modalities such as temperature.
5. Navigation Established the Perception-Action Loop at Minimal Cost
One of the first embodied-AI tasks to gain broad attention was object-goal navigation: place a robot in an unseen simulated environment without a prebuilt map, give it a target such as “a chair,” and ask it to find the object autonomously.
The key was not finding the chair itself but establishing a “perception-action loop.” The robot uses vision to choose an action; movement changes its coordinates and camera view; the new perception triggers the next action. Internet AI stops after identifying “this is a cat.” Closed-loop intelligence requires action to produce new feedback from the environment.
Manipulating objects can also change the environment, but it involves complicated physical contact. Navigation only requires moving the observer, making it closer to the research path that vision scholars could tackle at the time. Earlier one-step grasping often amounted to viewing a point cloud, predicting a grasp pose and stopping after execution—a form of open-loop 3D pattern recognition.
6. Navigation, Robot Learning and Foundation Models Eventually Converged
When CoRL was founded in 2017, it had already established the importance of learning for robotics, but it did not fly the Embodied AI flag. Robot learning could exclude vision and did not necessarily produce a closed loop; much of the grasping work at the time was still viewed as one-shot geometric inference.
Meta and other institutions later used the Habitat simulator and progressively richer tasks to make navigation the first flagship embodied-AI problem. Research then returned to closed-loop manipulation, connected navigation with manipulation, and finally used foundation models to open up the instruction space.
王鹤 believes several public statements completed the academic and industrial “naming” process. 李飞飞 identified embodied AI as one of computer vision’s future North Stars, while 黄仁勋 said the industry was approaching the next generation of AI—Embodied AI. The attention encouraged previously separate communities in vision, control and robot learning to absorb one another.
7. Young US Researchers Turned to Embodied AI Before ChatGPT
王鹤 recalls that around 2019, computer-vision and reinforcement-learning researchers preparing to enter the faculty market were already moving visibly toward robotics. Deepak, the co-founder of Skild AI, and Abhinav Gupta came through this academic path. Deepak’s early best-known work was curiosity-driven exploration.
Curiosity rewards address sparse feedback. If Mario receives 1 for finishing the level and -1 for dying, with 0 everywhere in between, reinforcement learning can barely explore its way to the endpoint. The model therefore rewards states where its prediction error is large, driving the character to keep discovering unexpected situations.
王鹤 crossed paths with Deepak while interning at FAIR Robotics Team in 2019 and also tried using reinforcement-learning methods such as A3C for tabletop object manipulation in simulation. The project was never published because the internship lasted only 3 months, but it strengthened his conviction about closed-loop manipulation.
8. His First AI Project in 2016 Already Contained a World Model
王鹤’s first AI project at Stanford was titled Learning a Generative Model of Multi-Step Human-Object Interaction from Videos. The team filmed tabletop manipulation videos, labeled action intervals and used an LSTM to learn action sequences and changes in object states.
The model had to connect “what the person did” with “how the world changed” and “what could be done next.” An empty bottle cannot pour water, while a full one can. Whether a hand is empty or holding something, and whether a cup lid is open or closed, changes the probability of the next action. This was effectively an early world model.
The output had two sides: multi-step 3D animation showing a person picking up a kettle, pouring water, putting it down and handing over a cup; and a smart cup that could drive itself. When it saw someone pick up a bottle containing water, it moved closer to receive the water. An empty bottle or a book triggered no response. The logic was learned from video rather than hand-written rules.
9. Moving from Semiconductors to AI Was a Deliberate Choice of Research Tempo
王鹤 entered Tsinghua through a physics competition and studied semiconductor and device physics as an undergraduate. His original workflow was to build a mathematical model by hand, fit it to experimental data and predict a new case. He says the structure resembles machine learning, except the fitter is a set of physical equations rather than a neural network.
After entering Stanford in 2015, he spent time in a clean room making germanium transistors, performing lithography and using ion-beam etching. An idea could take a month to fabricate and validate, while manual operations were full of randomness. Once he dropped a chip into hydrofluoric acid with tweezers; over-etching meant starting over.
“Thinking fast and working slowly” did not fit his mindset, so he left a physics track into which he had already invested 7 or 8 years. He joined his adviser Leo’s group and worked with a graphics postdoc on physical interaction. The first complex system was still built with Caffe, with no mature paradigm to copy.
10. The Scarce Research Skill Is Defining the Problem, Not Just Writing Code
Nine students in Leo’s group were competing for 2 PhD slots, including top computer-science students from Tsinghua and Shanghai Jiao Tong. 王鹤 admits that his coding horsepower was not the strongest. He stayed because he could formulate vague ideas as learnable problems.
He had to decide which object states needed to be extracted, how to segment the videos, how to train detectors and state classifiers for each category, and how to connect causality with a temporal network. The central diagram he drew when reporting to his adviser was “state—action—change world—state.”
The work took several years from its start in 2016 to publication. It was rejected twice at SIGGRAPH before finally being accepted at Eurographics in 2019, where it was nominated for best paper.
11. His Second Project Extended Instance Pose to Category-Level Understanding
王鹤 later acknowledged that the first project’s vision of learning everything from video and building a general world model has still not become the technology most capable of advancing deployment. The method that truly carried forward into 银河通用 came from the synthetic data used in his second project.
Traditional 6D pose estimation requires a 3D model, coordinate origin and xyz axes for a specific object in advance, then predicts its relative rotation and translation. That means every coffee mug must be modeled separately, making unseen instances of the same category impossible to handle.
Category-level pose estimation instead uses a canonical state shared by humans. Regardless of the pattern on the cup, people can say that a row of cups has its openings facing up and handles facing right. The model therefore only needs to know that an object belongs to the mug category, then describe a new instance relative to the category’s normal orientation.
12. IKEA Backgrounds and Digital Mugs Produced a Transferable Training Set
There was no ready-made real-world dataset for the category-level task. A PhD student could not photograph an infinite number of mugs and label their 6D poses one by one. 王鹤 therefore took an RGB-D camera to an IKEA in the Bay Area, photographed real tabletops, beds and other surfaces with depth, and extracted planes on which objects could be placed.
He rendered digital assets of mugs onto the real backgrounds: the background was real, the foreground synthetic, and the rendering process automatically supplied ground-truth poses. This produced hundreds of thousands of mixed-reality images, followed by mix-to-real testing.
The work became a 2019 CVPR oral presentation. Concepts such as NOCS (Normalized Object Coordinate Space) helped turn category-level pose estimation into a research field. By 2021, CVPR had a corresponding subfield; the paper was repeatedly reproduced by many researchers, while the dataset and synthetic data became classics in robot vision.
13. Going All In on Embodied AI in China Went Against the Consensus
In early 2020, 李开复 suggested that 王鹤, who specialized in 3D vision, LiDAR and point clouds, work on autonomous driving. 王鹤 replied that he wanted to study richer interactions between people and objects, hands and objects, and robots and objects, as well as home robots. The response came almost immediately: “Home robots are still 50 years away,” making the field unsuitable for investment at the time.
王鹤 did not dispute the timeline. He simply believed he could study the next wave from academia first. After returning to Peking University, senior professors advised him to keep at least half his effort in 3D vision, but he says he “just didn’t listen,” pushing the team toward robotics and closed-loop interaction.
His earliest domestic ally was his former schoolmate 卢策吾. 卢策吾 promoted the embodied-AI concept in the south and 王鹤 in the north; both participated in early papers and the construction of CCF terminology. 王鹤 wrote Exploring Embodied Intelligence in Yanyuan to build a narrative for the Chinese-language context.
He summarized the target as “a general intelligent agent capable of interacting in the physical world.” The name itself was unimportant. Learning interaction from human videos, perceiving object motion states and building home robots were different stages of the same research line.
14. Home Robots Are the Endgame, but China Offers More Room for Ownership
王鹤 chose home robots first because the space was largest. Autonomous driving solves only the problem of driving in the external world; homes need a general physical entity capable of doing housework and serving people. Google Everyday Robots was the reference image he most often used when explaining this vision.
He did not stay at Google or FAIR because his internship experience made researchers at large companies feel more like “screws in a machine,” with little ability to determine direction. As an Asian man, he also felt that the US offered limited room for advancement. Beijing was his hometown, and returning to China gave him a stronger sense of ownership.
At Peking University, he could immediately build his own team, combine 3D vision with reinforcement learning and win the inaugural ManiSkill Challenge world championship. Choosing a faculty position was not an escape from industry; it was a way to accumulate technology and autonomy for home robots before the market matured.
15. ChatGPT and PaLM-E Put “General-Purpose” Back into the Capital Model
Before ChatGPT, the biggest problem in embodied AI was that every skill required a new round of data collection. Once completed, it solved only one physical task and marginal costs did not fall. It could be a robotics method, but it was difficult to view it as revolutionary general-purpose technology.
PaLM-E introduced a model in which a vision-language foundation model served as the brain, small models were called to perform tasks and the system accepted open-ended instructions. 王鹤 emphasizes that under today’s strict definition, PaLM-E might not even qualify as truly embodied because it does not directly output actions. But the story was powerful enough to open investors’ imaginations.
If home robots truly become general-purpose, 王鹤’s market analogy is “more expensive than a car, larger in volume than smartphones,” potentially becoming the world’s largest industry. BAAI’s embodied AI center predates ChatGPT, but PaLM-E prompted investors to start searching for domestic embodied-AI teams in 2023.
16. 银河通用 Was Founded on the Premise That Hardware Capability Had to Be Filled In
When people first urged 王鹤 to start a company, he refused repeatedly. Management could be divided among co-founders; the real weakness was that he did not build hardware himself. At the time, the wheeled, legged and hand-equipped platforms he could buy could not even execute basic tasks reliably. “We didn’t even know how to talk about intelligence—execution itself was impossible.”
In early 2023, investors including IDG managing partner 李骁军 visited Peking University’s Jingyuan to discuss whether the direction made sense. 王鹤’s condition remained unchanged: someone had to fill the gaps in the robot body and mass-production capability. Software could not be built on a foundation where “all hardware is garbage.”
滕舟 brought experience in ABB robot mass production and desktop robots at a startup. Their combination created the division of labor across software, hardware and CEO responsibilities. 王鹤 calls the relationship a “spiral ascent”: the body determines whether intelligence can execute, while intelligence determines whether hardware creates value.
17. Wheeled Dual Arms Are a Deliberate Reduction in Complexity for Mass Production
银河 currently uses a wheeled base, dual 7-DOF arms and harmonic reducers. Most core components have been validated over the past decade. It has not immediately adopted Tesla-style legs, planetary roller screws or cable-driven mechanisms that have yet to reach mass production.
Unproven components must simultaneously confront yield, consistency, reliability and durability. Any single failure can stop the entire machine. 王鹤 believes that when intelligence is still the main bottleneck, adding aggressive hardware only slows shipment, repair and commercial iteration.
The wheeled base is about 60 centimeters in diameter and cannot pass through the tightest spaces in a home, but it is sufficient for supermarkets and factories. B2B customers ask only whether the robot works, whether it can work longer than a person and whether it performs better. They do not care whether it uses legs or wheels. “More precisely, it is pragmatic.”
18. Robots Must Clear 1,000, 10,000 and 100,000 Units in Sequence
王鹤 notes that the global industrial robot-arm industry generates roughly RMB100B in annual output, about the scale of a single leading automaker. Even leading commercial cleaning-robot companies sell only 10,000 to 20,000 units a year. Reaching 10,000 embodied robots would already be a phenomenon-level scale for the robotics industry.
A product cannot really be called mass-produced until it reaches several hundred thousand units, preferably 1M or more. Even after a century of automotive manufacturing, new models still expose yield and reliability problems. Humanoid robots that have never been mass-produced cannot skip these steps.
银河’s plan for 2025 is mass production at the 1,000-unit level. 王鹤 believes that after further convergence, the current wheeled platform can support at least 100,000 units and possibly 1M “without any problem.” Potential applications include supermarket loading and unloading, pharmacies, industrial box handling and palletizing, feeding and sorting.
Unmanned pharmacies require robots to work more than 20 hours a day. That is a different product from a demo in which a teleoperator steps in for 5 minutes before the robot is taken offstage. 银河 believes its hardware has undergone more complete iteration precisely because it has exposed and fixed problems under this intensity.
19. The First Bottleneck for VLMs and VLAs Is an Under-recorded World
王鹤 explains that LLMs are strong not only because of their architecture, but because internet text covers a large share of the language humans might produce. Public images and videos, by contrast, record only a small fraction of what the world’s 7B people see each day.
A VLM therefore cannot understand every image it receives and performs materially worse than an LLM. The action data needed by VLAs has only begun to be collected over the past 2 years, with even lower coverage. Generality cannot simply be extrapolated from language models to physical manipulation.
His judgment is explicit: a fully general VLA within 1 or 2 years is “at least from academia’s perspective, and from my own understanding, impossible.” The practical goal should be to develop intelligence around replicable applications, covering the full behavior set required by each application.
20. Constrain the Skill Set First, Then Deepen Generalization Across Objects and Environments
银河 is currently focused on a few atomic actions—moving, picking and placing—rather than trying to solve drilling, cutting and Rubik’s Cubes at the same time. But the same capability must handle different objects such as sparkling water, coffee and soft drinks, and deploy across 7-Eleven, FamilyMart, Lawson and different pharmacies.
The robots are already closed-loop systems. The VLM can understand whether a product on a shelf has fallen over, been placed crookedly or dropped, while the action system can pick it up from the floor or shelf and put it back. If asked to tear open a bag of gummies, it can understand the request and remove the bag, but it will not execute because tearing is untrained.
王鹤 calls these systems “behavior-constrained physical agents.” The constraint comes from manipulation data, not from a complete failure of visual understanding. The skill library will expand, but the small number of most important skills must first form a complete solution or a workable human-machine collaboration.
21. Productivity Comes from Repeatable Deployment
The same capability must survive changes in products, stores and operating conditions rather than remain tied to a single demonstration. A robot that works only in one carefully staged environment has not yet crossed into production.
The relevant test is whether a deployment can be copied with sharply lower marginal engineering cost. Rebuilding the model, workflow and delivery process for every site turns robotics into a labor business, regardless of how sophisticated the demo looks.
The path to scale is therefore to make a small number of high-value skills complete enough for routine operation, then extend them across customers and scenarios. The long-term goal can remain general-purpose, but commercial proof must arrive through repeatable deployments.
22. Complex Systems Push Market Share Toward a Few Leaders
张小珺 asked how demand for hundreds of thousands of units could be served if no single company could monopolize it. 王鹤’s answer was that robotics has a “very strong leader effect.” Industrial robot arms have the “Four Families,” while only a few commercial cleaning-robot companies, including Gaussian Robotics and Pudu, have reached 10,000-unit scale.
The reason is not brand marketing but system complexity. A cleaning robot can lose localization and enter an elevator, trip over construction barriers or flood a mall when its water tank malfunctions. Making one product work well can require nearly 2,000 people across hardware, software and services during peak periods.
Automobiles have a roughly 30M-unit market, mature supply chains and a century of industrial accumulation, yet manufacturers are still constantly eliminated. Embodied robots have neither mass production nor mature technology, making gaps in capital, talent and iteration data between leaders and the middle even harder to close.
23. Collecting Real Data at 10,000 Units Is a Bill No One Can Currently Afford
Autonomous-driving data arrived quickly because the car itself had useful functions. After buying a car, users continued generating road data for the manufacturer for free—arguably at “negative cost.” Robots without useful functions cannot be sold in the same way at 10,000 units and then left to customers for continued teleoperation.
王鹤 estimates that a full-size humanoid robot costs at least RMB100,000 to manufacture, meaning 10,000 units would require RMB1B for the bodies alone. If each robot operates 2 shifts with 2 people per shift, plus labeling and quality control, monthly labor expenses per unit would reach tens of thousands of yuan. The full system could cost hundreds of millions to RMB1B per month.
He believes 10,000 units would be a reasonable collection threshold if the industry relied entirely on real data, but “no one in the world can do this.” US companies often present their capabilities through video, making long-term open-site verification difficult. Videos alone cannot prove routine productivity.
Robot bodies are also being revised rapidly. A company may iterate through several versions in a year, with the dynamics and mechanical structure changing completely after a few months. Early machines cannot be treated like cars that remain stable for 10 years; simply accumulating 10,000 units does not mean obtaining data from the same distribution.
24. Synthetic Data Has Already Taken the Lead in 银河’s Training
银河’s public demonstrations of shelf VLA systems and unmanned pharmacies operating more than 20 hours a day are possible because the company does not rely entirely on real data. 王鹤 says real data accounts for roughly 1% or less of the training set, with most of the work handled by its in-house synthetic-data pipeline.
This advantage was not created simply by adding people. Counting from the NOCS work in 2017, he has spent roughly 8 years working on synthetic data and sim-to-real. Visual, physics and content gaps, as well as bias in robot-control modules, must each be examined and repaired.
王鹤 points to a sales loop in the industry: “I don’t believe in synthetic data. I tell you synthetic data is useless. You buy my machine, then you go shake it.” He believes some companies use simulation internally as well, but their pipelines are poorly built; after failure, they blame the approach itself.
25. Sim-to-Real Is Established; the Debate Is Now About Which Physical Processes Transfer
The response to “simulation physics is unrealistic, so sim-to-real is impossible” is that humanoid robots’ walking, running and jumping already rely on large-scale reinforcement learning in simulators. Locomotion capabilities from companies such as Unitree have shown that body state and contact dynamics can transfer.
Skeptics have therefore shifted the boundary to vision. 王鹤 cites 银河’s 2023 work showing that grasping transparent objects and glass fragments can be trained entirely with synthetic data and transferred to the real world. Even at the small-model stage, the visual gap can be narrowed systematically.
At the VLM-based VLA stage, he is even less convinced that rendered textures are the fundamental obstacle. A VLM that understands the real world can also understand Donald Duck, Mickey Mouse and film animation; the gap between physical rendering and live footage is far smaller than the gap between animation and reality. The real challenge is learning manipulation processes and their causality.
26. Synthetic Data Has Clear Boundaries; Deployment Feedback Handles the Final 2 Points
王鹤 acknowledges that the physics gap has not disappeared. Walking on the ground is relatively simple, while manipulation includes flexibility, compression, multipoint and multisurface contact, and fracture. 银河 judges current simulation sufficient for moving, picking and placing, and has begun testing dexterous hands, folding clothes and hanging clothes.
Tying shoelaces involves complex flexible contact and is difficult to simulate realistically enough at this stage. Tearing and cutting paper have been studied but may not deserve product priority. A supermarket solution does not need to tie shoelaces or tear open gummy bags; leaving those skills out for now does not prevent the system from becoming productive.
Synthetic data plus a small amount of real teleoperation can first push the success rate to roughly 98%. Real deployment feedback and reinforcement learning can then repair the remaining errors. 王鹤’s data flywheel is not about replacing simulation with real data, but using scarce field data to target the failure distribution.
27. The Real Product of the Embodied Era Is a Productivity-Grade Product
Research platforms are products too: a robot without a function can be sold to developers for continued research. But 王鹤 believes the healthy development of humanoid robotics now depends on moving from research platforms to “productivity-grade products.”
Productivity does not require a robot to handle every task in a space. Humans can cover areas the robot has not yet mastered. The standard is output per unit of time approaching human levels, with sufficient operating duration. If the robot is much slower and reduces efficiency, it is “backward productivity.”
银河’s long-term mission remains general-purpose, but it must “lay eggs along the way,” advancing hardware, skills and commercialization one step at a time. Otherwise the company will either regress into an academic institution or become a project business that cannot replicate at scale and has limited output.
28. Industry Verification Must Begin by Rejecting Hidden Teleoperation
王鹤 believes the US market is more tolerant of companies without product innovation, amplifying the embodied-AI bubble. Using Figure and PI as examples, he lays out a two-sided valuation logic: if robots can truly enter production lines reliably, long-term value could reach several trillion dollars; if there are only a dozen or 20 units and no routine operation, valuations in the tens of billions need to be reconsidered.
Some overseas demonstrations hide the operator backstage without clearly stating that the system is actively teleoperated. The more dangerous domestic disorder is: “I do not tell people I am teleoperating, but I am.” 银河’s first rule is: “Demonstrate publicly, and do not allow teleoperation.”
The second piece of evidence is sustained workload: how much work does the robot actually perform each day after entering the site, is there a continuous record, and can the platform certify it? 银河’s unmanned pharmacy processes hundreds of orders a day in routine operation. Videos and strategic agreements, by contrast, are becoming less persuasive.
29. Ten Thousand Units in 5 Years Is Both a Route Dispute and a Countdown for Industry Credibility
The robotics industry will have route disputes resembling Waymo versus Tesla for years. Some companies are structurally dependent on real data and lack an in-house synthetic-data pipeline; others keep investing in simulation. 王鹤 mentions NVIDIA’s environment-rendering platform built on NeRF, while internal simulation at some smaller autonomous-driving companies is “child’s play.”
Real-world collection will become an important support, but only once hardware is stable, production reaches scale and manufacturing costs are spread over enough units. Collection can also become more efficient through lower-cost locations such as Malaysia. If the industry has no company with annual revenue above RMB10B, it cannot sustainably spend tens of billions of yuan a year collecting data.
王鹤 requires leading Chinese and US companies to produce autonomous robot applications at the scale of “10,000 units in a single year” within 5 years. Otherwise the technology cycle will lengthen and capital enthusiasm will retreat: “Our field has been falsified; the bubble was all bubble.” The industry could even enter an ice age.
30. Demographic Constraints Turn Long-Term Idealism into an Immediate Responsibility
王鹤 believes China’s available labor force will continue shrinking every 5 years as the population ages and birth rates fall, potentially reaching “less than half of today’s level” over the next 10 to 20 years. Japan already faces labor shortages across age groups. If China loses 100M workers, even combining the populations of China, Japan and neighboring countries would not fill the gap.
The most dangerous promise is therefore: “Once you collect the data, you can train the model. Once you build the factory, you will have the skills. I sell you the robots, you collect the data, and tomorrow they become your employees.” If large-scale failures destroy the industry’s reputation, the damage will extend beyond company valuations. China could also lose a route for filling labor shortages in manufacturing and services.
Home robots will not become omnipotent overnight. 王鹤 expects products that do “light work, with heavier emphasis on companionship and entertainment” to appear first, then expand from light tasks such as retrieving drinks from the refrigerator to cleaning toilets and handling drains. Small-batch pilots may emerge within 3 to 5 years, while true generality still has decades of exploration ahead.
31. What 黄仁勋 Cares About Is Whether Synthetic Data Can Turn GPUs into Robot Infrastructure
王鹤 understands that he was seated next to 黄仁勋 because both NVIDIA and 银河通用 place heavy weight on synthetic data. GPUs can perform physics rendering and simulation while also training models. If the approach works, NVIDIA could support “half the sky” of embodied intelligence.
After several visits to 银河 by NVIDIA’s vice president of robotics and others, 王鹤 showed 黄仁勋 a series of projects completed entirely with synthetic data on his phone during dinner. Beyond the technology, he noticed that 黄仁勋 could handle spicy food, ate boiled fish with chili well and enthusiastically responded to a face-changing performance.
What left the stronger impression was 黄仁勋’s working style. Beyond his own table, he went table to table toasting and taking photos. 王鹤 described him as both approachable and a “model worker.” Even after building a company of that scale, the CEO still handled high-intensity communication personally.
32. Embodied AI’s Endgame Is Far Beyond VLA, but Value Does Not Require Full Generality
王鹤 warns that the industry’s understanding of human central motor control remains rudimentary. The cerebellum contains more neurons than the cerebrum, and motion generation and control are not divided as simply as public intuition suggests. Neither VLA nor a “large model wrapped around small models” has come close to exhausting the secrets of full-body coordination.
That means embodied AI cannot reach “full maturity and fulfillment” within 1 or 2 years, but specialized productivity can still generate economic value first. Using an annual salary of RMB200,000 per employee, he gives the example that a robot performing the work of 10,000 employees would represent RMB2B a year. Because robots may operate multiple shifts, the corresponding economic value could approach RMB10B without waiting for full generality.
A Brief History of Time taught him to question the essence of things and think from first principles. Romance of the Three Kingdoms led him to study strategy, character and organizational outcomes in turbulent periods. He most admires 曹操’s pragmatism and 诸葛亮’s idealism, hoping to “solve problems with 曹操’s methods” while preserving 诸葛亮’s original conviction.
That conviction also defines the social objective of robotics. If productivity is understood only as taking people’s jobs, the world will eventually oppose it. The sustainable direction is to address labor shortages while supporting healthy social development. In research, the focus should remain on 1 or 2 genuinely important breakthroughs, not on packaging “known-to-be-useless” inventions behind a higher paper count.