Pioneers Insight Method Research Author
Danfei Xu: Human Data, Behavior Cloning, Robotics’ GPT-3, Stanford, Full-Stack Robotics, EgoMimic, Teleoperation, UMI
Back to Episodes

Danfei Xu: Human Data, Behavior Cloning, Robotics’ GPT-3, Stanford, Full-Stack Robotics, EgoMimic, Teleoperation, UMI

Summary

  • Xu Danfei’s core thesis is not to use human data to top up teleoperation data, but to treat humans as another kind of robot and use first-person video, hand pose, VIO and other signals to recover as much as possible of an embodied agent’s inputs, goals and interaction patterns. He is explicit, however, that human video alone is unlikely to teach a robot the lowest-level control needed to generate joint movements and forces. His ultimate vision is: “I want to behavior clone a human”—a robot that can not only manipulate objects, but also act, communicate and interact like a person.
  • Behavior cloning was systematically underestimated. Xu saw at DeepMind in 2019 that “behavior cloning just works,” but RL was then the flagship research agenda and doing BC was “not politically correct.” After returning to Stanford, he spent 3 months building from a C++ control stack to ResNet-18, spatial softmax and an RNN, getting a Franka through roughly 30 seconds of oven manipulation. The model did not generalize, but it provided a critical “sign of life”: the bottleneck in robot learning is often not a new algorithm, but the complete system spanning teleoperation, control, latency, cameras and data distribution.
  • The scale target is 100 million hours, while the industry is still roughly 100x short. Xu estimates that human-level physical intelligence may require 100 million hours of human data. He has heard that the largest single-day collection amounted to roughly 100K–200K hours; even aggregating the highest-quality data on the market might yield only 1M–2M hours. The opportunity and risk are emerging together: “a few labs are laying tracks like crazy,” and capital is fueling the train, but data formats, sensor combinations and useful distributions have yet to converge.
  • UMI is the current sweet spot, while humanoid upper bodies are the long-term beneficiaries. UMI has people operate directly with a gripper that can be mounted on a robot, sacrificing five-finger dexterity for higher fidelity, scalability and a smaller end-effector gap; over time, it should become increasingly hard to distinguish from pure human data. Xu is unsure whether legs are necessary, but believes a humanoid upper body—with at least two arms and five-finger hands—is necessary. Human data gives humanoids a use case, while humanoid hardware lowers the human-to-robot transfer gap: both are outcomes of a hardware lottery and a data lottery.
  • The durable moat is more likely to be the integrated closed loop than any individual data-capture device. Project Aria-level VIO/SLAM, fast dexterous hands, robot-arm controllers, synchronization and calibration all require extensive engineering. The capture hardware itself may become commoditized, since frontier labs will also need to repeatedly teach vendors which data is useful. The key is the “data–training–deployment–evaluation” loop, and the team’s ability to understand how every segment of the data pipeline changes behavior.
  • Video may absorb most modalities, but it cannot solve control by itself. Xu currently ranks video, hand pose and language highest; whole-body pose and tactile sit in the middle, while audio and smell rank lower. A wrist-mounted RGB camera might even replace part of the role of tactile sensing and occluded-state estimation. But robots still have to learn for themselves how to generate joint movements and forces; under the current paradigm of treating human data as teleoperation data, precise state estimation remains important.
  • The robot equivalent of a GPT-3 moment is not human-level perfection, but a usability inflection point in generalization: in any setting, for anything a person can do, a robot should succeed roughly 40%–50% of the time. The strongest counterarguments would be simulation proving far more scalable than expected, the second or third human-to-robot gap resisting zero-shot or few-shot transfer, or non-humanoid embodiments proliferating first. On the modeling side, the system would need long context and a “new language,” because natural-language planning remains “too far away” from precise manipulation.

Deep dive

1. Interest Was Not a Personality Detail but a Hard Constraint on Xu Danfei’s Path

  • Xu Danfei was born in Taiyuan, Shanxi, and moved to Shanghai between sixth grade and the start of middle school. In junior and senior high, he bought microcontrollers, small cars and tutorials on Taobao and tinkered on his own. School was mainly a social venue and “where I spend most my day,” not a source of positive feedback through grades.

  • He describes himself as extremely interest-driven: for things he likes, he invests “anywhere between fifty percent, hundred percent”; for things he does not, “zero percent effort.” It is not that he cannot tolerate hard work. He simply cannot persuade himself to do something whose meaning he cannot see at the time.

  • Looking back, he realized that he was good at handling high-uncertainty, open-ended problems. He was of course anxious at the time, but the counter-pressure to return to the gaokao track was even stronger. Solving one problem after another on his own gradually convinced him that “many things are doable by yourself.”

2. A DIY U.S. College Application Turned Nonconformity into an Advantage for the First Time

  • After visiting the U.S. in his first year of high school, Xu decided to live differently: “there’s no way to turn back from there.” Starting around the middle of his second year, he spent less time at school and gathered information on the ACT, TOEFL and applications through QQ and the old CUUS forums. He used no agent and handled every document and process himself.

  • When 何泰然 asked whether leaving the gaokao track created pressure, Xu’s answer was unequivocal: leaving the mainstream was exciting. Only 2 or 3 students in his class at Shanghai Shangnan High School were likely applying to U.S. colleges, but whether something was mainstream was never a decision variable for him.

  • The application process was not smooth from start to finish. He first entered a liberal arts college such as Dickinson College, but saw no problem with the choice: more room for exploration and closer relationships with professors gave him an entry point into research. His parents did not understand the U.S. education system; they simply paid for exams when he needed to take them.

3. Twenty-Plus Cold Calls Bought His First Robotics Research Role

  • As a college freshman, Xu reached a threshold in “the desire to do research” and directly called more than 20 companies with robotics in their names, including Boston Dynamics. SynTouch Robotics ultimately took the bet: “Then come. We won’t pay you either, but come.”

  • Jeremy Fishel, the mentor who answered the call and later became a longtime friend, may not have fully understood Xu’s English, which was still limited in spoken conversation. But he saw the passion and initiative. Xu’s own summary was not a motivational slogan, but a familiar decision test: “what to lose.”

  • At SynTouch, he used a BioTac tactile sensor for Bayesian tactile sensing. The sensor simultaneously measured pressure, the speed at which a material conducted heat and high-frequency vibration. It was heated to near human skin temperature, sampled vibration at roughly 2,000 Hz and used those signals to infer the material being touched.

  • It was also his first intensive exposure to Shadow Hand. By the time many others later heard about the dexterous hand, he had already “broken several.” He went through roughly 7 or 8 fingers in a single summer, confirming his hardware affinity: “actually sit next to robot and see it move and break it and fix it.”

4. CMU’s Self-Driving Project Made Him Fall for Complete Systems Rather Than Ready-Made Datasets

  • When CMU RISS recruited its first cohort, a researcher working on autonomous driving replied that Xu’s background “might” be a fit. Xu did not wait for the process to continue; he drove roughly 4 hours to Pittsburgh, knocked on the office door and spoke with the researcher for half an hour. The outcome followed the same logic: if it was possible, “make it happen.”

  • The project tackled city-wide RGB localization without GPS. A modified Jeep carried 6 cameras and 6 computers. Xu and another researcher drove around Pittsburgh collecting images roughly 2 days a week—“we just drove around Pittsburgh”—rather than downloading a dataset and working only on the algorithm.

  • Time alignment across the cameras and multiple servers became a critical bottleneck. He once spent an entire night in the lab fixing the network timeline. AlexNet had already appeared, but CMU’s Robotics Institute was still dominated by classical robotics. What truly excited him was that the cameras, network, vehicle and algorithm all had to work together.

5. Stanford’s Attraction Was Precisely the Absence of a Ready-Made Answer

  • When applying for a PhD, Xu wavered among CMU, U-Dub and Stanford until April 14, one day before the deadline. Stanford was nearly a robotics desert at the time, but the other paths were too predetermined. He had a vague sense that “there’s something bigger that I can do” there.

  • His first rotation involved end-to-end 3D reconstruction. By the second, he had begun working on egocentric VR and human-data capture. He bought an Oculus DK and mounted Leap Motion on it to capture hand movements—a setup nearly identical to his later initial VR-plus-camera prototype.

  • After completing work on scene-graph generation with 菲菲 and Yuke Zhu, 菲菲 asked whether he wanted to continue. Xu replied: “No, I want to do robotics.” The absence of hardware and an established team was not an obstacle. He was energized by the ability to lead everything himself: “I hate other people telling me what to do.”

6. Early Robot Learning Was Caught Between Structural Priors and Reinforcement Learning

  • Around 2016–2017, robot learning had roughly 2 mainstream camps. One treated robotics as a vision problem, focusing on grasping and one-step pick-and-place. The other, inspired by AlphaGo, bet on online/offline RL and motor babbling. Long-horizon tasks on real robots remained rare.

  • The prevailing bias was binary: either write the prior to be sufficiently intelligent, using structure, dynamics or a neurosymbolic program, or let the robot learn directly from the environment. Supervised learning was seen as “not good enough, not scalable,” and was even treated as something embarrassing.

  • Xu initially wanted to build one-shot imitation learning—“demonstration in, action out.” After discovering that OpenAI had pursued a similar direction, he added more structure and developed Neural Task Programming. That decision established his early-PhD preference for compositional structure.

7. He Did Not Abandon Compositionality, but He Stopped Insisting on Hand-Coded Structure

  • Xu later revised the method, not the problem: compositionality can be the objective without being written into the model structure in advance. The essence of task and motion planning is to break an enormous, intractable search space into local spaces and then handle the connections among them.

  • 何泰然 challenged the approach on the grounds that human-designed decomposition, while offering built-in generalization, also creates a ceiling on the tasks that can be solved. Xu believes task and motion planning may still work in constrained settings such as factories, and notes that production systems may use behavior trees or finite-state machines.

  • He still retains the vision of generative task and motion planning. The approach only starts “getting there”—approaching a Bitter Lesson version—when a trajectory plan can be represented directly as video and generated in the space by a deep generative model.

8. DeepMind Let Him See BC Work Firsthand—and See How the Research Agenda Suppressed the Evidence

  • During a 2019 internship at DeepMind, Xu studied generative imitation and GAIL. Training GAIL at one point consumed roughly 10% of the entire lab’s compute. In hindsight, he saw the setup as seductive but prone to detours: it assumes demonstrations are limited and rewards are unavailable, but that the agent can interact extensively with the environment. On real robots, interaction is precisely what is hard, while demonstrations are relatively obtainable.

  • The more important observation came from the infrastructure. The team teleoperated Sawyer with a SpaceMouse and had a large volume of high-quality data, including failures and reward labels. Once less-than-optimal trajectories were filtered out, simply doing behavior cloning was highly competitive: “I saw with my own eyes behavior cloning just works.”

  • The issue was not that nobody saw the result. Reinforcement learning was the flagship product and flagship research agenda. Xu put it sharply: BC was “forced down,” and “it’s not politically correct to do behavior cloning.” That was the mood across the community at the time.

  • He could not see a credible path from RL to scaling complex behaviors such as cooking a meal. Once BC was demonstrably working, the central question shifted from “why not use a more advanced algorithm?” to “why not do teleoperation, data and supervised learning properly?”

9. The Franka Project Showed That the First Challenge in Behavior Cloning Was Systems Engineering

  • The form of behavior cloning is simple: a person controls a robot with an iPhone, VR controller or another interface while the system records camera images and control commands, treating the pair as supervised-learning x and y and training a policy that directly outputs actions.

  • Xu and J immediately clicked. J had previously worked on Arm Farm, collecting data with 4 or 5 robots and iPhones. But the team was influenced by BC shame and preferred to pursue offline-first learning that could work in theory but was difficult to make work in practice. The 2 concluded that “offline is not needed” and turned to making BC work properly.

  • In roughly 3 months, they built a force-control teleoperation stack for the Franka Panda from the C++ control layer all the way to the model. A wrist camera, ResNet-18 encoder, spatial softmax and RNN history were all “why not” choices, but they worked together inside a complete system.

  • The most striking demo was a roughly 30-second, 1×-speed long-horizon task: take a plate out of an oven, move objects, close the oven and put the plate back. It “did not generalize at all,” but few people had previously seen comparable behavior on a real robot, making it an important “sign of life.”

10. Even an Effective Paper Needed a Novelty Wrapper, and COVID Cut Off the Follow-Up Experiments

  • Even though both authors believed “BC is going to work,” they worried that simply proving BC effective would not clear peer review, so they added a more publishable story around the system. Xu did not reduce academic taste to doing only what works, but believes the community must dismantle the notion that effective methods deserve shame.

  • The project convinced him that camera placement, controller response, camera latency, robot hardware and data distribution can each determine success or failure. “If you don’t pay attention to any one of these, you’re definitely going to have problems.” Strong systems therefore should not be judged only by isolated algorithmic novelty.

  • The team had planned to spend a year producing 4 or 5 papers along this line, with the next step being to make DAgger genuinely serve BC. The COVID shutdown took away their access to robots. The paper had little external impact; many people simply responded that “a Stanford student did it with a better system.”

  • The line of work that became ALOHA did not drive a broader paradigm shift until around 2023. Xu admits that, in hindsight, he should have pushed further. But there were only 2 of them, and they did not know whether to prioritize the model, action space or data. He was already in his fifth or sixth year of the PhD and beginning to look for faculty jobs, leaving no way to build a large research agenda covering multiple hypotheses.

11. Three Internships Told Him Which Problems Were Boring and Which Organizations Were Intolerable

  • His time at Zoox convinced Xu not to work on driving again. Autonomous driving had deteriorated from the complete robotics system he liked in 2013 into a clearly partitioned 3D vision pipeline: perception, planning, control and simulation teams each optimized their own metric, and cross-team communication was no longer necessary to solve the problem.

  • He does not think general-purpose robotics will mature the same way, at least along the path he believes in, because the problem cannot be solved by splitting it into isolated modules: “everyone needs to know everything.” That became the cultural core of full-stack robotics.

  • DeepMind’s shift from exploratory research toward top-down management pushed him toward a faculty career: only entrepreneurship or academia seemed likely to preserve the freedom to “do things you’re interested in.” By 2026, however, he also acknowledges that resource intensity has changed the answer. Without existing resources, he might lean toward industry, because some experiments genuinely cannot be run in academia.

12. Robot Learning Changes the Method Stack, Not the Problems Robots Need to Solve

  • Xu defines robot learning as leaving the problems—planning, precise manipulation and whole-body control—unchanged while replacing as much traditional modeling as possible with data-driven learning. Whole-body control, for example, moves from hand-coded dynamics plus optimization toward reinforcement learning.

  • Traditional manipulation explicitly writes different dynamics before and after contact, sometimes producing a mixed-integer program. The new paradigm cares more about whether the robot outputs the correct action and whether the object reaches its goal; the intervening physics model no longer needs to exist explicitly.

  • He believes the importance of the model and algorithm has been overestimated in recent years, while the system comprising hardware, control, data and deployment has been underestimated. Full stack does not mean building everything in-house. It means knowing which details can change the final outcome and understanding them deeply enough.

13. Human Data Is Not Auxiliary Data; It Rewrites the Training Paradigm for Robots

  • Current data falls broadly into 3 categories: teleoperation across various robots, synthetic data generated by physics engines and real-world non-robot data. The last includes unconstrained data such as YouTube videos, as well as purpose-collected first-person or third-person human-operation data.

  • Xu’s skepticism toward teleoperation comes from a specific fact: even a slight change in a robot’s low-level controller gain can make the same control command produce a completely different response. If the robot-to-robot gap is already large, the human-to-robot gap may not be fundamentally unbridgeable.

  • If human movement can be converted into usable actions and perception into policy inputs, human trajectories could directly serve as robot trajectories. His key analogy is that “a human is just another robot”; the reverse can also be said: “a robot is just another human.”

  • That leads to a broader sensor vision: “If you put enough sensors on every person, you can actually turn a person into a robot.” But this does not mean a single video can recover complete control. He is explicit that human video is particularly poor at directly teaching a robot how to generate joint movements and forces.

14. EgoMimic’s Zero-to-One Work Was First About Reliably Turning a Person into a Measurable System

  • EgoMimic was Xu’s first or second project after joining the faculty. A collaborator believed strongly that first-person data might be the most scalable. Xu worried that starting with hardware would make the project drag on, but after roughly a week of discussion they still decided: “let’s just do a system.”

  • The first version followed his 2015 idea: it initially used an Oculus plus Leap Motion, then switched to Quest and Orion, bringing head localization, hand tracking and video into a common coordinate frame. Calibration was extremely unstable, and the gaps among the 3 devices kept the team struggling for weeks before the system worked reliably.

  • The turning point was Project Aria. The glasses provided first-person RGB, head localization and hand-state estimation in one package, covering the core requirements. Even so, the team spent roughly 4 or 5 months getting data collection to work end to end. Meta’s related teams even created a Project P0 for them to handle every support request.

  • The initial focus was to make both the first-person data-collection system and a robot body with a more human-like embodiment work in practice. Scalability, data distribution and transfer remained questions for later research.

15. To Use Human-Like Data, He First Built a More Human-Like Robot Himself

  • There was no off-the-shelf platform combining 2 arms, human-like shoulders and torso joints with a camera near eye level. Xu bought aluminum components from Vention and personally designed and assembled an experimental robot with 2 arms plus shoulders and a torso.

  • As an assistant professor, he spent time in the lab “tightening screws” whenever he was not teaching. When the original gripper kept breaking, he designed and 3D-printed a replacement at home, then one day handed it to his students and said: “Use this.”

  • His conviction took shape only in the middle and later stages of the project. Teleoperation already creates controller and embodiment gaps among robots; if perception and action can be normalized, there is no impassable boundary between humans and robots. That shift in view came from building and operating the system.

16. Human Video Can Teach Three Layers of Knowledge, but the Lowest-Level Joint Control Still Belongs to the Robot

  • Xu breaks embodied interaction into 3 layers. First, how the world should change. Second, where and how the embodiment should interact with the world to cause that change. Third, how the embodiment should generate joint movements and forces to produce the required motion.

  • Human data naturally contains the first layer: the goal and resulting world change in picking up a cup or pushing something 5 centimeters do not depend on the executor. If a robot has near-human five-finger hands, workspace and force capacity, the second layer can also transfer—for example, which part of the cup to grasp or where to push.

  • The third layer is the hardest to obtain directly from video. Throwing a ball is the representative example: the robot must generate enough force at every joint to send the ball along a specific trajectory. Human EMG signals and motor commands do not have a meaningful one-to-one mapping either.

  • That is why he places precise manipulation and physical common sense ahead of language planning. Task planning is relatively straightforward. The real difficulty is making a robot demonstrate physical common sense through manipulation—the enormous gap between the symbolic layer and the physical layer.

17. First-Person Data Is Not the Cheapest, but It Offers a Viable Fidelity–Scalability Trade-Off Today

  • 何泰然’s objection is worth preserving: third-person video is far more abundant on YouTube, and humans can quickly learn from a video of someone replacing a car filter or assembling furniture. Why insist on first-person data? Xu summarizes the conflict as fidelity versus scalability.

  • Third-person data is easy to download, but its viewpoint, scene and behavior distributions differ too much from a robot’s inputs. Searching all of YouTube for clips that align with the robot’s distribution, then normalizing them, may leave very little usable data. Purpose-collecting egocentric data is expensive, but more consistent.

  • He acknowledges that precise state estimation may matter less over the long run, but that corresponds to a different paradigm: human data tells the robot how to do something, while the robot already knows how to reach every state. In the current approach of treating human data as teleoperation data, more precise state estimation is required.

18. VIO and SLAM Determine Whether Human Data Can Acquire Genuine Action Labels

  • The difference between the 2 current paradigms is this. In one, human data only tells the robot how to complete a task, while the robot already knows how to reach each state. In the other, the human is treated as a robot and human data is converted into teleoperation data for behavior cloning. The latter must capture as much of the full action and state as possible.

  • Hand-pose estimation must be placed in the same coordinate frame as head localization before hand movement can be converted into actions in world coordinates. That makes the glasses’ own position, VIO and SLAM part of the core stack.

  • Xu believes the state of the art in sub-centimeter SLAM sits at AR/VR companies such as Meta and Apple. Project Aria can even calibrate lens relationships online as they deform with temperature. The gap between an academic city-scale pipeline and that capability is “like comparing a VLM with SIFT.” But this remains an engineering problem that money and time can grind down, not a theoretical barrier requiring the next generation of GPT.

19. Tactile Matters, but Force Is the More Fundamental Physical Variable

  • Xu calls tactile the second-most important data modality, while remaining unsure how much value a wrist-mounted RGB camera could absorb. A first-person camera is often occluded by the hand itself; a wrist camera can directly see contact between hand and object and may use the mature computer-vision industry to solve much of the occlusion problem.

  • His theory is that an agent is fundamentally a force-exertion engine. Its role in the world is to apply force and change the world’s state. Tactile is the most direct—but imperfect—measurement of force, not the ultimate target itself.

  • Force and tactile should be present on both sides of the interface. The robot must know the force state acting on itself, and it must generate and regulate the force it applies to the world. The difficulty is that optical, pressure, resistive and magnetic sensors all use fundamentally different representations, with no standardized unit analogous to RGB.

  • Asked to give up modalities, he put video, hand pose and language first; whole-body pose and tactile toward the back; and audio and smell first to go. He still retained uncertainty about tactile, because “we don’t know how to use it yet” does not mean it is unimportant.

20. UMI Trades Away Human-Hand Freedom for High-Fidelity Data Close to the Robot’s Own Interface

  • UMI, or Universal Interface, can be understood as temporarily replacing a human hand with a robot gripper. The person directly holds the gripper and operates it while the system records its 3D pose and opening state. The same gripper can then be mounted on the robot, keeping the end effector and collision-model environment gap small.

  • 何泰然 describes it as “forcibly degrading the human hand into a gripper.” The cost is five-finger dexterous manipulation. The benefits are high-quality state estimation and directly deployable actions, with better scalability than traditional teleoperation.

  • Xu still classifies UMI as human data because the person directly interacts with the environment, can feel force through the gripper and can control it precisely. There is no robot controller in the middle translating intent a second time. A gap in the robot’s workspace remains, but pure human video faces the same problem.

  • On the product of fidelity and scalability, UMI is an obvious sweet spot today: it offers more direct action and state information than pure video and is easier to scale than teleoperation.

21. UMI and Pure Human Data Will Converge, Shifting the Real Bottleneck to Robot Actuators

  • Xu expects UMI and human data to become difficult to distinguish over time. Generalist’s system, for example, simply places 2 fingertips and a camera on the human hand. Remove the fingertips and it becomes pure human-hand data; keep them and it looks more like gripper data. The difference is the degree of fidelity and transfer.

  • From the data perspective, the 2 categories will ultimately share a great deal. The more robotic the human collection device, the more direct the transfer. The more native human capability it preserves, the richer the data—but the more the final result depends on whether the robot hand has comparable capabilities.

  • What many teams have truly failed to get right is the low-level controller and whole-robot integration. The robot hand’s closing, arm movement and response speed need to match the pace of human execution. Xu believes Generalist’s system is especially strong on this point, while most failures are misdiagnosed as data or model problems.

  • For five-finger transfer, his vision is specific: a dexterous hand slightly smaller than the 22-DOF Shadow Hand, but faster, paired with a higher-performance arm and fully integrated low-level actuators, could achieve UMI-level transfer. The current bottleneck lies more in the robot hardware chain than in the capture side itself.

22. Human Data and Humanoids Are Not Substitutes; Each Gives the Other Value

  • Xu is unsure whether 2 legs are necessary, but believes a robot needs at least a humanoid upper body with 2 arms and 5 fingers. The lower body should at minimum have a holonomic base. The closer the morphology is to a human’s, the easier it is to turn existing human-operation data into a training asset.

  • Asked whether this is a hardware lottery or a data lottery, his answer is both. A humanoid is not necessarily required by the human environment itself, but human data gives the humanoid a clear use case. Conversely, without a human-like embodiment, human-to-robot transfer becomes materially harder.

  • 何泰然 asked whether this would lock robots at human performance. Xu invoked the null space: a higher-DOF system can imitate a lower-DOF human by reducing its behavior, just as a 20-DOF hand can simulate a single gripper. Human data provides a starting point; it does not require the final robot to have only human-like joints.

  • The area that may be genuinely overlooked is human–robot interaction data. Teleoperation can collect object manipulation, but it does not naturally generate interaction, cooperation and reactions between people. If HRI is also treated as a learning problem, that data can only come from human interaction.

23. “Behavior Clone a Human” Is the Upper Bound, Not the Complete Training Regime After the Starting Point

  • Given sufficient data, compute and hardware, Xu’s upper bound is a human-like robot indistinguishable from a person—one that can “talk like a human, behave like a human, interact like a human,” not merely manipulate objects but also interact with other people.

  • He compresses the vision into one line: “I want to behavior clone a human.” In one sense, today’s LLMs first behavior-cloned human language, then became superhuman coders through supervised fine-tuning and reinforcement learning. Robots have not yet completed even that first step.

  • Human data therefore need not permanently constrain superhuman embodiment. Once robots pass a basic-capability inflection point analogous to GPT-2 to GPT-3, RL or other fine-tuning can unlock actions and morphologies that exceed the human data. That is the more sensible sequence.

24. 100 Million Hours Is the Target Scale; the Scarcer Asset Is a Continuously Improving Performance Curve

  • Xu estimates that the vision may require roughly 100 million hours of human data. He has heard of a maximum daily collection of roughly 100K–200K hours. Even a well-funded company aggregating the market’s high-quality data might obtain only 1M–2M hours, still roughly 100x short.

  • He does not see a 100x gap as insurmountable: “it can be done.” 何泰然 then raised data transmission, storage, management, filtering and other infrastructure. Xu warned that human data still lacks an MP4-like standard representation; even the modalities and sensors to collect have not converged, so blindly scaling volume could waste resources.

  • Xu has roughly counted more than 100 companies now collecting this type of data. Initially, he expected Meta Glass, Ray-Ban Glass and other AI wearables to generate first-person video incidentally over the next 5 years. The robot-learning boom has instead made dedicated collection the leading path much earlier.

  • The risk is equally concrete: “a few labs are laying tracks like crazy,” while capital keeps “adding fuel and coal” to the high-speed train. But data standards and useful distributions remain unsettled.

25. Incidental Human Behavior Contains More Physical Intelligence and Is Harder to Clone Correctly

  • Xu emphasizes incidental human data, because much of physical intelligence never appears in a task specification: opening a door with a foot, closing a drawer with an elbow or picking something up without knocking over the object beside it all emerge from natural responses to everyday constraints.

  • If a data company tells collectors to optimize only for task completion, people will actively reduce complex and rich interactions. The ideal framework may sit between the extremes: a small amount of intentional task data to prove a capability, a large volume of daily life to provide the distribution and then key-data filtering, much as in autonomous driving.

  • The challenge is that “behavior clone an agent” requires interpreting context. If a person suddenly checks the time while cooking because they do not know the time, should a robot with its own clock clone that behavior? How should meaningless postures caused by neck discomfort be filtered? These are the hard questions in understanding why a behavior occurred.

26. Data-Capture Hardware Will Become a Commodity; the Research-to-Deployment Loop Can Form the Moat

  • Xu believes the capture pipeline is more likely to become commoditized over time. Frontier labs themselves do not fully know what data they need and must continually teach vendors what to collect. Vendors also serve other customers, making it difficult to keep devices, processes and useful distributions proprietary forever.

  • If he were designing a company, he would favor a Tesla-style vertically integrated model that “owns everything.” But industry reality is more likely to produce a structure in which a frontier lab drives several data vendors: the research institution defines the problem, while suppliers handle collection at scale.

  • The open-data vision around EgoMimic is to use the credibility of researchers and collaborators to consolidate vendor data, run quality checks and prove through research which data works, making the vendors’ data itself more trustworthy.

  • 何泰然 points out the paradox: “your success equals your failure.” If the route truly requires 100 million hours, intense capitalization, consolidation and closed operations may be unavoidable. The thing that ultimately binds the market may not be a particular piece of hardware, but commercial interests: if buyers of the data do not share it, the industry will recreate autonomous driving’s data barriers.

27. If the Human-Data Thesis Fails, the Counterevidence Will Come from Three Foundational Assumptions

  • The first possibility is that Xu has materially underestimated the scalability of simulation-based learning. If physics engines and synthetic data can cover the real world, expensive human collection will no longer be the only scalable source.

  • The second is that the second and third gaps between humans and robots are too large to support zero-shot or few-shot transfer.

  • The third is that non-humanoid robots proliferate first: 3-finger or 3-pronged systems, highly capable robots or robots that can rotate indefinitely could make humanoid data increasingly difficult to align.

28. Full Stack Is Not About Building Everything In-House; It Is About Owning the Causal Loop

  • Teams can buy hardware, data and tools, but evaluation and the training loop must remain internal. Researchers need to personally inspect post-training, pre-training, filtering, collection quality and deployment outcomes, and understand how each type of data changes policy behavior.

  • The ideal academic loop has the same researcher participate in data collection, model training, deployment and evaluation. If vendor data becomes a black box, the team may not even know whether synchronization is reliable. At 100 million hours, that invisibility becomes a major challenge.

  • The clearest inductive bias human data provides for modeling is long context. Without enough history, the model cannot explain why a person took a particular action; it simply learns a broad distribution in which the behavior could be “A or B.”

  • Both System 1 and System 2 should be learned from human data, but the interface need not be natural language. Xu believes “we need a new language”—a shared latent space for planning and manipulation learned from scratch. LLM-generated language plans are still too far from action. Betty the Crow can combine tools to solve a new problem when confronted with several of them for the first time; robots remain “too far” from zero-shot generalization even when they can overfit through RL.

29. Research Taste Comes from Understanding the Gradient of a Direction, Not Chasing an Existing Trajectory

  • Xu’s advising style swings between hands-on and hands-off, but the lab culture is clear: care about the whole thing; no work is beneath anyone. When a robotics module breaks and nobody knows how to solder it, he will solder it himself. He wants students willing to become full-stack researchers.

  • When selecting students, he admits that his bias is that they “cannot hate hardware.” Young researchers should work on complete projects early to build a holistic understanding of systems, then decide whether to pursue open-ended research, solve a single problem or make industry practice more rigorous.

  • In his view, doing a robotics PhD in 2026 is harder than it was 10 years ago. FOMO is stronger, working outside the mainstream can make students feel left out and industry “currently has a winner,” leaving less patience for developing independent taste. The counterpoint is that resources and coding tools are far more abundant: “if you choose correctly, it’s much easier to do something.”

  • He advises young researchers to watch good researchers and ask why they do something, not what they did, learning the gradient rather than the trajectory. His personal ambition is simply to help create a robot GPT-3 moment: a robot that can perform anything a person can do in any setting with roughly 40%–50% success. His time capsule for the future is: “Do what you want to do… being outside the main group sometimes doesn’t have such bad consequences. What’s to lose?”