Pioneers Insight Method Research Author
E206 | The Robot GPT-3 Moment Is Near: Accelerating Evolution of Open-Source Embodied AI Models
Back to Episodes

E206 | The Robot GPT-3 Moment Is Near: Accelerating Evolution of Open-Source Embodied AI Models

Summary

  • The investable shift in robotics models in 2025 is that unified foundation models are rewriting the objective from “perfecting one task” to average success across hundreds or thousands of tasks. π0 used folding clothes to validate complex, deformable-object manipulation, while π0.5 moved into previously unseen homes; Wall-OSS is focused on generalization and long-horizon tasks. 王昊 calls this an “exponential effect” in applications: models are no longer limited to one motion, but can autonomously interleave steps, plan, and complete compound workflows such as clearing a dining table.

  • The biggest enemy of generalization is not conventional visual noise, but the physical world’s infinite, impossible-to-label long tail. A wrinkle in a tablecloth, glare from a transparent object, or a cup placed slightly off-balance can cause small errors to snowball through a long-horizon task; π0.5 outperforms π0 in new environments, but Kay cautions that it is not a model that “can be dropped into any new environment and perform tasks well.” Ultimately, models need more than memorized scenes: they need commonsense physics, spatial reasoning, and the ability to model future processes.

  • The core data moat for robotics companies is not simply accumulating hours, but building a deployment loop that produces data that is simultaneously abundant, high-quality, and fast to collect. 王昊 estimates that leading companies still have tens of thousands to hundreds of thousands of hours of data, while Wall-OSS used tens of thousands of hours of real-world data; Kay offers a deliberately rough “million-hour” thought experiment, equivalent to roughly 100 years of human physical experience. Real interaction data remains irreplaceable: synthetic data mainly helps close visual-distribution gaps, while human video is better at conveying action intent and task planning.

  • Without a reproducible real-robot leaderboard, it is far harder to verify model leadership in robotics than in LLMs. Kay says reality is often less a clear win by one model than “a fight among mediocre players”: authors double as evaluators, while hardware condition, environmental long tails, and task selection can all change the result. For investors, robot competitions, clothes-folding demos, and grasping demos prove only local capabilities; more credible signals are data efficiency on the same hardware, cross-environment generalization, and success rates in open-ended settings.

  • End-to-end training versus hierarchical control remains unsettled, but 王昊 proposes a pragmatic path: train jointly, split at deployment. He advocates aligning language, vision, and action in a shared space with joint backpropagation, then putting slow reasoning in the cloud and fast control at the edge; Kay believes the industry has not even securely reached the “GPT-2 Moment,” so architecture choices should remain subordinate to data. π0 and π0.5’s 50Hz is not a magic number: the key is an Action Chunk of roughly one second and 50 steps per output, which gives models with limited memory and observation a more coherent short-term plan.

  • The two guests fundamentally disagree on where the industry stands, but 王昊 sees Scaling Law as the clearest direction forward. Kay judges current capabilities to be below GPT-2; 王昊 believes scaling gains are already proven, that robotics is now at the GPT-2 stage, and that it can reach GPT-3 in “one to two years.” That forecast assumes continued expansion of data, models, and embodied infrastructure—not a sudden breakthrough from one isolated algorithm.

  • The commercial battleground is choosing scenarios that can generate revenue while supplying open-ended data for general-purpose models. 王昊 expects robots to cook simple dishes and wash dishes in semi-structured kitchens within 2-3 years, and potentially enter open kitchens in about 5 years while tolerating errors through human collaboration; Kay gives a more conservative 5-10 years, comparing the path to early robot vacuums, where users learned to work within the product’s limitations. China is more likely to advance on two tracks—foundation models and scenario deployment—and actual deployment volume, hardware cost reductions, and the data flywheel will determine who can fund the journey to a viable product.

Deep dive

1. Unified Foundation Models Are Ending the “One System, One Task” Era

  • Kay looks back on 7-8 years of robotics research: the field could previously make one difficult task work extremely well, but the full solution was “very hard and expensive to replicate conveniently for a new problem.” The biggest change over the past 2-3 years is that model generality has finally been validated enough for researchers to discuss generalization and improvements in overall performance.

  • 王昊 places the inflection point around 2023. Before then, the goal was to perfect individual tasks; now, one unified foundation model can learn and execute hundreds or thousands of tasks at once, with the optimization target shifting to average success across all tasks. The “exponential effect in applications” comes directly from this change in the objective function.

  • Moving into robotics also changed 王昊’s understanding of data. Language and multimodal models rely on static internet data, while continuous physical processes captured at second-level or centimeter-level resolution remain poorly documented. Embodied AI is therefore not simply about scaling models, but about advancing hardware, data, and models as an integrated stack—the “next frontier of AI.”

2. π0 to π0.5 Takes Generalization from the Same Task to Unseen Homes

  • π0 chose folding clothes as its representative task in 2024. Kay explains that one extra wrinkle or a slightly different fold angle can constitute a new state for a robot; adding a long sequence of decisions about what to fold first and what to fold next turns a trivial human chore into a problem the robotics field has studied for 10-20 years.

  • π0.5 was subsequently deployed on a mobile robot and placed in homes absent from its training data. It showed some promise in unfamiliar homes, much as a person still knows how to pick things up in someone else’s house, but Kay stresses that the result was “not perfect.” Examples such as “cups of different shapes” are progressive tests of grasping generalization and do not all mean π0.5 has already completed those tasks.

  • Grasping itself has layers of generalization. Changing the background or position of the same cup is one layer; handling a similar-looking object made of a different material and requiring a different strategy is another; the ultimate goal is for a robot hearing “pick up the cup” to find and grasp a completely different cup in an unfamiliar home.

  • There is no established formula for data collection. Teams can only keep asking: “We collected data from 3 more houses—did that help?” They may even wonder whether the approach is wrong if 30 homes bring no improvement. The eventual result showed that π0.5 is not necessarily much better than π0 in the same environment, but is stronger in new environments.

3. Long-Horizon Tasks Bind Generalization, Planning, and Complex-Case Handling into One Problem

  • 王昊 argues that any robot task deployed in a real setting becomes a long and complex workflow. Clearing a dining table involves hard utensils, liquids and food scraps, irregular trash, and flexible towels, while also requiring the robot to decide where each object belongs and avoid spills.

  • These tasks have no fixed “do A, then do B” sequence. A robot may need to alternate between wiping the table, emptying scraps, organizing utensils, and folding towels while handling unexpected situations. Humans cannot easily draw the boundaries of every subtask in advance, so the model must make end-to-end decisions, plan in real time, and execute in a closed loop.

  • The training variables were primarily homes, with other settings also covered, including setting and clearing tables and organizing bathrooms and rooms. 王昊 says the robot’s manipulation and algorithmic capabilities in these long-sequence closed-loop tasks gave the team “a major confidence boost,” though the program did not provide a unified success rate.

4. The Physical Long Tail Turns Small Errors at Each Step into Final Failure

  • 王昊 defines the genuinely difficult part as corner cases that cannot be enumerated during training. Lighting and conventional visual errors can be mitigated with sensors, compute, generative models, and data augmentation, but “the physical world itself has infinite possibilities,” and many variations cannot be labeled in advance.

  • His examples: one extra wrinkle in a tablecloth can make a cup unstable, while a single reflection from a transparent object can happen to interfere with the camera. Humans adjust instantly through intuition and experience; highly data-driven models may not truly “feel” what the change means.

  • Long-horizon tasks amplify the problem. A single-step error may be small, but it can be “snowballed” into a much larger error that ultimately causes the entire task to fail. Models therefore need not only to recognize the scene in front of them, but also to possess commonsense physics, spatial understanding, spatial reasoning, and an internal understanding of physical processes.

  • 王昊’s conclusion is not that the answer is simply to “go collect data in the real world like crazy.” Real-robot data, human video, and other sources must be organized into a larger, higher-quality, more diverse data system. The bottleneck is therefore data engineering and the data pipeline, including collection costs, filtering, cleaning, and task distribution.

5. Without a Real-Robot Leaderboard, Models Remain Trapped in “Authors as Evaluators”

  • Kay points out that language models can announce their rankings on leaderboards, while robotics has failed for decades to establish an objective, fair, and reproducible real-robot benchmark. How many environments and long-tail cases a task should include, and whether hardware maintenance details affect the outcome, can all destroy comparability.

  • Papers can usually only have their authors design the tests and decide whether a new algorithm is better. In the best case, a strong model is “obviously better”; in reality, it is more often “a fight among mediocre players,” where a change in model, data, or algorithm appears effective but the source of the improvement is hard to identify.

  • Soccer, racing, and task demos in robot competitions are similarly difficult to use as universal standards. 王昊’s practical alternative is to compare open-source models on the same robot body and task by data requirements, generalization, and reasoning ability, or place different approaches in the same real-world application and test them under randomized conditions.

6. Beyond the Data Trilemma, Hardware Maintenance Still Eats Research Velocity

  • Kay summarizes the data dilemma in 2025 as a continuing trade-off between quality and quantity. High-value data requires careful design, collection, and cleaning; once teams focus on detail, scaling becomes difficult. Yet models need data that is “more, better, and faster.”

  • Real-robot research also faces a basic operational hurdle: a researcher may spend “the whole day doing nothing but repairing the hand and tightening screws.” Robotics still lacks a broadly accepted, stable, easy-to-use hardware platform analogous to software that runs after download, and the industry continues to debate what such a platform should look like.

  • When π0 was released, PI estimated that its data volume already exceeded the total used by earlier related Google research, even though PI was still a young company. Kay also observes that real-robot data does not have a fixed value forever: as teams accumulate collection experience, mature processes gradually lower costs.

  • 王昊 estimates that leading robotics companies still have roughly tens of thousands to hundreds of thousands of hours of data, far below the data scale of GPT-4-level language models. Costs cannot be compared directly across companies: quality, diversity, hardware and labor prices, as well as the ability to build, reset, and automate operations in scenarios, all change the value of each data hour.

7. “A Million Hours” Is a Starting Point for Exploration, Not a Mechanical Copy of Human Lifetimes

  • Kay calls her view a “bold idea”: using a 100-year lifespan as a rough estimate, a human accumulates about 1M hours of physical experience, while there appears to be no publicly known robot dataset with 1M hours. She suspects that once the industry reaches that scale, “perhaps we will only then begin the exploration that follows.”

  • The figure does not mean one robot must train for 100 years. Once robots are widely deployed, many machines can copy and share experience, so 1M hours might be collected in a matter of days. That could be the scaling inflection point for robotics data.

  • 王昊 cautions that comparing robots with babies is unfair. Humans have written physical experience and response strategies into their genes through long evolution, while continuously optimizing their hardware. His analogy is that biology “uses hardware wherever intelligence is unnecessary”; E. coli needs only chemical and temperature sensing to adapt to its environment.

  • Babies also do not spend 10,000 consecutive hours learning to grasp. They play with blocks and perform other tasks, then may naturally master a previously unfamiliar motion when they return to it a month later. Multitask learning extracts shared physical structure, which is the core logic behind pretraining robots on diverse tasks, reducing the data needed for new tasks, and PI’s exploration of cross-embodiment transfer.

8. Real-Robot, Synthetic, and Human-Video Data Each Solve Only Part of the Problem

  • Real-robot data is the most expensive, but remains the main source of high-fidelity physical interaction. It requires hardware, facilities, and operators, and collection speed is constrained. Cost reductions could come from cheaper robot bodies and sensor-equipped wearable devices, rather than requiring every datapoint to be collected in the final product form.

  • 王昊 believes generative models are useful for narrowing the gap between visual and real-world distributions, but are unlikely to synthesize data containing genuine physical interaction processes. Backgrounds and appearances can be augmented; contact, deformation, and dynamics still depend primarily on real-world collection.

  • Human video is massive, diverse, and relatively cheap. For now, it is better at teaching robots action intent, high-level semantic understanding, and task planning—and that planning can be learned “through video rather than through language.” Converting video knowledge into precise robot actions remains difficult.

  • π0.5 added web data to supplement outside knowledge and “general-purpose and commonsense” understanding, but Kay does not treat synthetic data’s value as settled. 王昊 points to Genie 3 as evidence that large amounts of high-quality data can be obtained from the internet, chiefly game environments, while generating video and performing some action control. It could become an environment for interaction training, but remains simpler than the real world.

9. Wall-OSS Attempts to Put Chain-of-Thought, Spatial Reasoning, and Action Generation into One Model

  • 自变量 trained Wall-OSS on tens of thousands of hours of real-world data and expanded an existing vision-language model so that one unified framework could perform “chain-of-thought” reasoning and generate actions. 王昊 highlights visual understanding, spatial reasoning, multilingual instruction following, and relatively high action-generation accuracy.

  • The model is primarily focused on generalization and long-horizon tasks. Solving long sequences necessarily means handling changing scenes, unseen objects, and multiple failure modes. Long-horizon capability is therefore not an independent metric, but a combined test of generalization, reasoning, planning, and execution.

  • Both companies view open source as a way to lower the research barrier. Kay says true open source requires restructuring code, testing it, and confirming that the community can run it; the payoff is seeing the model appear on robots the team never imagined. 王昊 believes it enables universities and small companies to build applications on foundation models, shifting recognition from winning the paper race toward engineering contributions.

10. End-to-End Training Does Not Mean Stuffing a 100B-Parameter Model into a Robot Unchanged

  • 王昊 firmly supports data-driven end-to-end training. Language, vision, and action should be represented and aligned in the same space; imposing manual hierarchy creates information loss. Models can reach tens of billions or even hundreds of billions of parameters while learning understanding, reasoning, and action generation together.

  • Deployment introduces physical constraints that require the system to be split. Tens of billions to hundreds of billions of parameters cannot all run at high frequency on the edge, so slow task reasoning can run in the cloud while fast physical processes remain on-device. The action component can also be distilled and compressed, while language and visual reasoning remain in a larger model.

  • Asked about a “brain and cerebellum” System 2/System 1 architecture, 王昊 says training may not cleanly divide parameters into slow and fast systems, but the overall system should receive one unified gradient update. Hierarchical optimization belongs primarily to inference and deployment.

  • Kay is more open on architecture. She believes robot models have not yet stably reached the “GPT-2 Moment,” so the key conviction for now should be data and data-driven algorithms. Whether reasoning and control are separated or combined should serve the data, hardware, and final performance—not become a battle between preset schools.

11. VLA Is Converging Research Paths, but Has Not Settled the “GPT-2 Moment” Debate

  • Kay remembers that imitation learning was not mainstream when she worked on it in 2018. Robotics research was fragmented across humanoids, vehicles, hands, PR2, and other platforms. As VLA—vision-language-action models—became popular, researchers using different methods began trying large models together, bringing upper bodies, lower bodies, and multiple embodiments into one general framework.

  • On the industry’s stage, the two guests do not seek consensus. Kay judges that current models “have not reached anything like a GPT-2 Moment”; 王昊 believes the GPT-1-style proof of concept is already behind the industry, with capability gains from increasing parameters and data now visible. In his view, robotics is currently at the GPT-2 stage.

  • 王昊 further predicts that robotics can reach GPT-3-level capability in “one to two years.” His basis is that language models had to pass through a period of fragmented exploration, while robotics has already seen the reliability of Scaling Law. The clear next step is to expand data, model scale, and real embodied infrastructure, then wait for emergent capabilities.

12. The US-China Divide Is a Combination of Compute, Manufacturing, and Survival Conditions—not “General-Purpose versus Vertical”

  • 王昊 characterizes the US path as top-down: use leading chips and massive compute clusters to first explore extremely large models approaching AGI. China’s chip constraints force greater efficiency, but he rejects the idea that China should focus only on “small and precise” systems, because it also has mobile application scenarios and a complete hardware supply chain.

  • His argument is that “you need a large and general foundation before you can have small and precise development.” China is better suited to a two-track approach that combines top-down and bottom-up execution: iterate on general-purpose models while entering scenarios that can test generalization, forming a commercial loop and data flywheel early.

  • Kay explains the difference through startup history. Chinese companies tend to start from user needs and survival, while the manufacturing base is particularly well suited to rapidly refining vertical products such as lawn-mowing robots. The current wave of US general-purpose robotics companies emerged mainly around 2024, partly because OpenAI’s success in general-purpose language models delivered an industry-wide “shock and rethink.”

  • Covariant Robotics provides an example. Kay says it initially also wanted to bring machine-learning robots into general real-world applications, but its success in logistics led outsiders to remember it as a vertical company. PI’s goal, she says, is to build general, data-driven systems, so it is careful about avoiding short-term commercial projects.

13. The Value of 50Hz Lies in Planning One Second Ahead, Not in the Frequency Itself

  • Asked whether the brain or the hand is harder, Kay answers, “both are very hard,” but believes algorithms can compensate for imprecise general-purpose hardware. During her PhD, she used ordinary hardware to achieve sub-millimeter precision, picking up a small ball with chopsticks and moving it through the air. Her principle is: “A black cat or a white cat—if it catches mice, it’s a good cat.”

  • She does not consider 50Hz inherently high-frequency. Earlier research often used 100Hz to above 200Hz, Berkeley grasping work has used roughly 5Hz to 20Hz, and robot low-level controllers can reach 1,000Hz. PI chose 50Hz mainly because the model generates a plan of roughly one second and 50 steps at a time.

  • This is the practical value of Action Chunking. Current models have limited observation, historical memory, and scene awareness, so making frame-by-frame decisions can produce hesitation. Planning one second into the future instead creates more coherent behavior. Kay says the approach, which originated in work at Stanford, improved learning from human demonstrations and has gradually become standard practice.

  • 王昊 believes control frequency is not the core issue. An excessively high frequency can sometimes indicate that the model is poor at predicting the future and can only perform short actions based on the immediate scene. Models with tens of billions to hundreds of billions of parameters already face significant pressure running at roughly 50Hz to 60Hz on the edge; if the vision model better understands motion and future states, the required frequency may actually fall.

14. Tactile Sensors Are Not Scarce; Force Data for Pretraining Is

  • Many current manipulation tasks rely mainly on vision. Identifying a tabletop and placing a cup accurately does not necessarily require touch. During grasping, an end-effector camera can observe deformation, rebound, and other cues to infer whether the object is secure and how much force is being applied—effectively “using vision to sense part of the force.”

  • 王昊 lists plenty of existing tactile hardware: multidirectional fingertip force sensors, electronic skin on arms and palms, temperature sensors, six-axis force sensors at the end effector, and joint torque sensors. The main reason touch has not entered foundation models is therefore not that the sensors are entirely immature.

  • The larger constraint is the historical distribution of data. Human-generated data is overwhelmingly visual, while touch and force have not been recorded at the same scale. Since vision already supports many critical tasks, the industry can continue accumulating force data and gradually fill the gap; tactile sensing is not yet an “absolute bottleneck” for robot control.

15. Home Robots Will Enter the Market First with Errors Users Can Tolerate

  • 王昊 views household service robots as a “perfect Turing test” for robotics. Chopping vegetables requires precise motion and force control; grease and ingredients require rich sensing; recipes and appliance instructions require long-horizon planning; fragile objects and unexpected events require robust response. The task combines nearly every challenge in embodied AI.

  • His timeline: within 2-3 years, models and robots could cook simple dishes and wash dishes in semi-structured kitchens; in about 5 years, they may enter fully open kitchens. Zero errors will not be necessary. If overall success rates are high enough, and robots can ask humans for help and collaborate with them, household products may become viable.

  • Kay offers a more cautious estimate of 5-10 years, comparing the path to early robot vacuums. A product does not need to be perfect at launch if users understand “what it can and cannot do” and errors remain within an acceptable range. A robot that occasionally fails can still create commercial value.

  • 王昊 summarizes the startup strategy as both “looking up at the stars” and “keeping one’s feet on the ground”: prioritize open settings such as public services and eldercare that are close to the general-purpose goal, while avoiding closed tasks. The more robots are deployed across diverse environments, the stronger the feedback and data loop; model improvements and hardware cost reductions will determine whether household demand can ultimately support the companies building them.