Pioneers Insight Method Research Author
148: Tashi Zhihang’s Chen Yilun: Three Dawns and the First Gate for Embodied AI
Back to Episodes

148: Tashi Zhihang’s Chen Yilun: Three Dawns and the First Gate for Embodied AI

Summary

  • The key bottleneck for embodied AI today is not whether VLA or World Model wins, but getting through the “data wall” first. Chen Yilun uses autonomous driving to calibrate the scale: he saw a clear inflection after accumulating roughly 10,000 hours; product-grade systems typically require 100,000 to 1 million hours, while more capable embodied systems may need at least 10 million hours. Tashi Zhihang has collected roughly 100,000 hours and says the unit cost of wearable data capture is at least 2 orders of magnitude lower than teleoperation.
  • The three dawns Chen Yilun sees are RL unlocking locomotion, GPT unlocking task planning, and end-to-end systems unlocking the full perception-to-action loop. Commercialization must still clear three walls in sequence: data, compute, and interaction. The first phase delivers the steepest performance gains through data; the second uses compute to “digest” that data; only in the third do World Models, reinforcement learning, and problem-specific creative post-training come into play.
  • His conviction in end-to-end systems came from a high-pressure, high-risk autonomous-driving experiment that compressed roughly 2 million lines of rules into about 300 lines of network code and deployed roughly 100 cars to collect human-driving data. After data accumulated to several thousand and eventually about 10,000 hours, the network moved smoothly through mixed traffic in urban villages where rules were nearly impossible to write; that was when he concluded that “AI can do planning,” and came to view L2 as a problem for which the key had already been found.
  • Tashi Zhihang is betting that robots should have their own foundation model, rather than simply growing an “action head” on top of a VLM. AWE (AI World Engine) allocates its main neurons to representing time, space, force, and how the world evolves after an action; Chen’s view is that robots operate in a contact-based physical world, where image-and-language data alone cannot teach a system to manipulate cloth, cables, and other objects continually changed by its actions.
  • World Model and VLA are neither synonyms nor simple competitors on his roadmap. If autonomous driving’s hardest current problem is negotiating with vehicles and pedestrians, a closed-system World Model and reinforcement learning should take it to L3 and L4; once the problem becomes an open world of road signs, unfamiliar obstacles, and puddles, language and VLA should extend the system toward L4 and L5—“every technology exists to solve the problem it is meant to solve.”
  • He expects 2026 to bring a “data explosion” and a growing number of demos, while conservatively forecasting that 2027 will definitely deliver effects comparable to end-to-end autonomous driving in 2021. The more credible signal of leadership will not be surface-level partnership announcements, but whether a team can build a genuine “value pool” in verticals such as industrial manufacturing; Tashi Zhihang’s first targets are wiring-harness handling, connector insertion, flexible assembly, and other productivity problems that traditional automation has failed to solve.
  • Chen Yilun’s view of Chinese teams is exceptionally forceful: in the embodied era, “American entrepreneurs will not be competitors to Chinese entrepreneurs—not at all.” His rationale is not merely supply-chain cost, but the repeated coordination required across hardware, sensors, robot bodies, scenarios, data, and algorithms. His long-term filter is whether a company can define itself and create real value that customers are willing to share, rather than chase superficial demos or partnerships.

Deep dive

1. Robotics Has Been Chen Yilun’s Long-Term Throughline

  • Chen Yilun earned an early admission through physics competitions, then moved between electrical engineering at Tsinghua and a US PhD in machine learning before turning toward the electromechanical world; he envied his roommate for building things that could “move,” while his own work at the time was entirely algorithmic.
  • In 2007, he watched Boston Dynamics’ hydraulic robot dog stay stable on ice and was “completely stunned.” After finishing his PhD, he did not take the mainstream AI route; instead, he learned about motors and servo control at his first company, and even personally led work on precision servo valves and hydraulic products.
  • That seemingly indirect path reflected a durable judgment: the algorithms of the time could only produce simple robots, “not the kind of robot I wanted,” so he kept waiting for the moment when AI was truly ready.

2. 2 Million Lines of Rules Pushed Autonomous Driving Toward End-to-End

  • By 2020, the team’s autonomous-driving system had roughly 2 million lines of code and could handle complex urban roads, but problems were arriving far faster than fixes: “we simply couldn’t keep piling it on.”
  • Perception was not Chen Yilun’s biggest headache: with data and ground truth, it could be optimized continuously. Planning and control, by contrast, formed a closed-loop AI problem—every action changed the next moment’s environment. After a cut-in, for example, the other driver might yield or accelerate to compete.
  • He and Ding Wenchao, among others, tried having a neural network plan trajectories directly. They eventually trained the network with roughly 300 lines of code in an early two-stage end-to-end setup; this was not an imitation of Tesla, whose AI Day at the time focused mainly on reconstructing 3D environments from vision and had not yet made end-to-end planning and control an industry trend.

3. Roughly 100 Cars and Urban-Village Testing Produced His GPT Moment

  • The experiment required large-scale collection of human driving. The team put roughly half its fleet—about 100 cars—into the project, while Ding Wenchao taught drivers every day what counted as “good driving.” Chen recalls that they proceeded under enormous pressure and risk, and only saw something different after accumulating about 10,000 hours.
  • The early data produced no obvious results. After several thousand hours, the network began to learn; as the data continued to grow, its capabilities kept improving.
  • The team chose unstructured urban-village traffic, where cars and people competed for space, and insisted on using as little post-processing as possible. When the network moved smoothly through the scene, Chen realized for the first time, with real force: “AI can do planning.”

4. He Left During the Industry Boom Because L2 Had Become an Engineering Problem

  • In 2022, as advanced driver assistance was scaling rapidly, Chen Yilun concluded that the key to L2 had already been found and that the remaining work was primarily engineering; the existing organization had the capability to turn it into a first-rate product.
  • He did not yet know how AI could solve L4, but end-to-end planning had proved that autonomous driving was merely one subproblem of general-purpose robotics. When he left, he told his managers and colleagues, “I’m going to work on robots next,” prompting mostly disbelief.
  • He did not start a company immediately, because his definition of entrepreneurship was “building a company that provides excellent products and solves customer problems,” and neither the market nor the technology was ready in 2022. He later joined Tsinghua’s AIR, where he continued to assess the timing of a robotics venture in an environment closely connected to industry; work with the founding team began last year, and the company was formally established on February 5 this year.

5. The First Dawn: RL Has Unlocked Locomotion

  • Chen Yilun believes MIT’s 2019 Mini Cheetah made its main contribution by opening up the complete hardware-software stack for quadruped robots. The ETH team subsequently found the “golden key”: using RL and neural networks to control full-body motion directly.
  • The approach depends on 2 pieces of infrastructure: highly concurrent simulators capable of generating massive training volumes, and hardware companies that reduce the digital-to-physical, or sim-to-real, gap through design.
  • That is why companies with strong hardware capabilities are often the ones that make robots dance, perform martial arts, and walk most fluidly. In Chen’s view, the core mechanics of locomotion are now understood; what remains is mainly time and iteration. “No one will worry about locomotion being a problem anymore.”

6. GPT and End-to-End Complete the Other 2 Dawns

  • GPT solves task planning. A trip to the Oriental Pearl Tower can be broken down by Google Maps or Baidu Maps into turn-by-turn navigation, but robots do not have a shared task map; large models are particularly good at decomposing a natural-language request into a sequence of steps.
  • The third dawn comes from end-to-end systems: robots must connect high-dimensional sensors and low-dimensional commands to the final action. Traditional methods require a large number of specialized modules and rules, while Chen has already seen in autonomous driving that neural networks can carry planning through the entire chain.
  • Embodied AI today looks more like autonomous driving in 2019: the logic appears sound, but the decisive effect has not fully arrived. The difference is that he has already seen the outcome once in a subproblem, and the past year of embodied experiments has “not shown anything beyond expectations.”

7. End-to-End Is Now a Consensus, but the Industry Still Lacks an Aha Moment

  • Early autonomous driving saw results first, followed by an industry-wide rush. Embodied AI has gone the other way: “end-to-end is something every single person is talking about,” but genuinely astonishing results remain broadly absent.
  • Chen Yilun defines end-to-end broadly: use neural networks to solve as many problems as possible and obtain performance through data that the previous generation of technology could not match. Internally, the approach may use imitation learning or reinforcement learning.
  • VLA only specifies that video, language, and action enter the same network. World Model may refer to generating future video from any viewpoint, or to an interactive model that takes state and action as inputs and predicts the next state. The terminology is not standardized; the key question is still which task the system actually solves.

8. Data, Compute, and Interaction Are 3 Walls That Must Be Crossed in Sequence

  • The first is the data wall: only sufficient data allows the network to grow sufficiently complex. GPT has internet-scale text by default, while autonomous driving must acquire data proactively through products and business models.
  • The second is often called the algorithm-and-compute wall, but Chen Yilun believes the real competition is usually compute. As data grows, network architectures may actually become simpler: “a heavy sword needs no edge; the greatest skill appears effortless.” Only a return to fundamentals can withstand the force of big data.
  • The third wall appears once scaling approaches its limits. Pretraining and compute are no longer enough, so teams must use post-training and problem-specific methods to truly “break through” difficult tasks.
  • Intelligence in the first phase resembles observation triggering associations with training fragments; it does not seriously reason about why an action worked. Once interaction begins, the model must predict how actions change the world, and only then can success rates potentially move to another level.

9. World Model First Solves Traffic Negotiation; VLA Later Solves the Open World

  • Chen Yilun believes autonomous driving’s hardest current problem is interaction over shared road space: whether to change lanes aggressively or conservatively, whether to brake or yield. Every decision changes how nearby vehicles and people respond.
  • Solving this game requires a driving World Model capable of rolling out parallel worlds, combined with reinforcement learning to select actions. It can even be a closed system, because the model only needs to capture interactions among vehicles, pedestrians, and the ego vehicle.
  • He believes this step will probably be sufficient for L3 and L4. The move from L4 to L5 comes later, when navigation information is insufficient and the system must read road signs and handle unfamiliar obstacles and standing water in an open world; at that point, language and VLA become more important.

10. GPT’s Greatest Invention Was Not the Architecture but Next-Token Prediction

  • Cheng Manqi suggested that the Transformer might have been the inflection point for large models in 2017. Chen Yilun’s answer was that OpenAI’s greater invention was finding next-token prediction: having the network constantly perform “fill-in-the-blank” exercises somehow led to today’s GPT capabilities.
  • He recalls an early blog post by Andrej Karpathy showing that a modest RNN could write poetry and code simply by predicting the next word. “No one was talking about GPT at the time,” but Chen’s first thought after reading it was whether the method could be used for autonomous driving.
  • The Transformer’s value lies in its simplicity, computational efficiency, and low error rate, allowing it to withstand massive data. It may not dominate on small datasets, but it matches Chen’s experience: “The more complex the task and the larger the dataset, the simpler the network architecture becomes.”

11. AI for the Physical World Needs a More Fundamental Representation Than RGB

  • Another “awesome” invention in autonomous driving was BEV. Whether the system is end-to-end or VLA, it is difficult to avoid first reconstructing video into space and then deriving planning from that spatial representation; without this layer, the network is closer to mechanically memorizing, “when I see this video, move this way.”
  • Chen Yilun extends the same idea to robots. The same object can produce an infinite number of viewpoints and pixels, but the more fundamental representation is where and when it occupies space, how objects relate to one another, and how they change under force.
  • Autonomous driving and robotics are therefore both forms of “physical-world AI.” What deserves the most neural representation is not texture and color, but the basic physical quantities of time, space, contact, force, and torque.

12. Robots Are Not Collision-Avoidance Systems; They Must Understand Contact

  • Autonomous driving is fundamentally a non-collision system, concerned mainly with how objects are arranged in space. Robots are contact systems: beyond x, y, and z, they must also understand force and torque.
  • Moving a hard block after grasping it is relatively simple because the object barely changes the task in return. When handling cloth, embroidering with a needle, or organizing soft cables, every action changes the object, so the model must continuously predict feedback.
  • Chen Yilun’s intuitive Turing test for embodied AI is to have a robot dress someone or put a hat on them, then see whether an observer watching only the motion can distinguish a human from a machine. Reaching that bar in a sufficiently complex constrained scenario, with a scalable method, would already constitute a major breakthrough.

13. Today’s Models Mostly Interpolate; Truly Efficient Learning Still Awaits a Value Function

  • Chen Yilun explains many apparent “emergent” abilities as combinations and interpolations of fragments in the training data. A user’s question can usually be traced back to relevant fragments, and the network recombines them into something that looks new, but is not created from nothing.
  • What genuinely generalizes first is therefore methodology. An AI coding model need not read all of Shakespeare and the Hundred Schools of Thought; it can solve code problems using the same method. An autonomous-driving system trained in China may not work in Japan or India, but it can extend there by adding local data.
  • Multi-task capability comes from continuously expanding the data and model, while deployment does not necessarily require every capability to be packed into one infinitely large model. “Methodology supporting generalization” is more realistic than learning every task in a single zero-shot leap.
  • He cited a direction Ilya discussed in a talk: humans may have inherited an excellent value function that lets them actively seek out and absorb data efficiently. If that can be cracked, learning efficiency will rise dramatically. For now, the effective approach remains a powerful data generator paired with a data fitter.

14. Embodied AI May Need to Start With Tens of Millions of Hours of Data

  • Drawing on autonomous-driving experience, Chen Yilun estimates that roughly 10,000 hours were enough for him to see a first inflection; a usable system generally requires 100,000 to 1 million hours, while leading companies often have more than 1 million hours.
  • If embodied AI is an order of magnitude more capable than autonomous driving, its data requirements may also be at least an order of magnitude higher. That puts his starting point at 10 million hours, rather than the small datasets common in laboratories.
  • This was the first data problem Tashi Zhihang tackled before starting the company. Models and tasks can continue to evolve, but if the data source cannot grow continuously at low cost, the team is not even qualified to compete on the later walls of compute and interaction.

15. Internet Videos Are Insufficient in Quantity, Quality, and Task Relevance

  • Autonomous-driving teams once scraped dashcam footage and YouTube travel videos, only to find that the true scale quickly topped out and quality was worse than expected. More fatal was the mismatch: a travel video from Tibet could not solve the problem of getting through a particular intersection.
  • Using Berkeley’s BDD100K dataset as an example, Chen Yilun said many excellent researchers eventually “gave up one after another” after completing this kind of dataset, because isolated static video was difficult to map onto the specific problem where the system was failing.
  • Responding to embodied companies that claim they can learn from vast quantities of video, he said explicitly that his team had “drawn a cross” through that route. The most valuable data must match a specific problem; nominal scale alone is not enough.

16. Simulation Works Only for Sufficiently Simple Subproblems

  • The team once had more than 30 excellent engineers rebuild large sections of a city, complete with rendered rain, snow, and standing water. The result looked impressive but did little for the core planning task, because it mainly improved visual realism, while perception was not the hardest problem.
  • Traditional physics simulation also has limits. Cutting a cable into small segments and tuning springs and Young’s modulus is still essentially approximating the world with rules; AI is already being used downstream while the data source remains hand-tuned rules.
  • The exception he accepts is locomotion. The biomechanics of a multi-joint body are simple enough and the environment can temporarily be ignored, allowing simulators to generate useful data efficiently. Even then, it took the industry about 2 years to make the approach work well.

17. A Human-Centric Data Engine Starts by Recording People, Not Teleoperating Machines

  • Chen Yilun divides the data sources for large-scale AI into people and the world. The most direct path is often “start with people, then move from people to the world”: language data is typed by people, while autonomous-driving data records people driving.
  • The true way autonomous driving obtained data was to record human driving behavior at the lowest possible cost. Given today’s visual capabilities, he would even consider putting a small dashcam in Didi vehicles instead of deploying 8 cameras and lidar.
  • The same applies to robots: people wear sensors to “see what people see and feel what people feel.” Once the data trains an AI, the robot extracts human experience from the model. Teleoperation merely transfers movements in real time; AI transfers skills.
  • He insists that data must capture both real scenarios and real actions. A purpose-built autonomous-driving test track cannot replace driving in a real city, and an artificially designed robotic environment cannot represent real production. The person collecting the data must also already know how to complete the task correctly.

18. One Glove Must Record Position, Pose, Fingers, and Force

  • Tashi Zhihang’s wearable system combines first-person vision with a lightweight glove, aiming to fully record the hand as an end effector: its position, pose, every finger’s state, and the force applied through tactile sensing.
  • A camera alone cannot localize the hand. When folding a quilt, the hand disappears beneath the fabric; reaching in darkness or dealing with occlusion creates the same failure mode. The team is therefore developing its own hardware and algorithms to keep the signal stable and accurate.
  • VR hand tracking still falls short in reconstruction quality and occlusion handling. Film motion capture only needs to convey broad movement and may not require millimeter-level operational accuracy. Chen has not found a complete glove designed specifically for embodied-AI training that can simply be purchased off the shelf.
  • Large-scale compute can be used in an auto-labeling-like data-processing stage, keeping wearable devices from being constrained by power consumption and cost. The glove itself also contains a small chip running a downsized model, moving as much computation as possible to the edge.

19. Wearable Capture Cuts Teleoperation Costs by at Least 100x

  • Teleoperation requires pushing a robot into the field and also suffers from latency, slow movements, and low success rates. A factory must first remove a worker from the line, while a café would be continuously disrupted by both the robot and its operator.
  • Chen Yilun’s rule-of-thumb comparison is that wearable capture can produce 10 usable examples out of 10, while teleoperation may produce at most 1. At the same data scale, Tashi Zhihang’s cost is “at least” 2 orders of magnitude lower—roughly 1% of the cost.
  • The team began formally scaling up in the second half of the year and is currently at roughly 100,000 hours. The main cost is not the glove or payments to data collectors, but compute.
  • The team first brought hardware costs down to a scalable level before expanding collection. Chen expects data to grow by many multiples next year and expects capture equipment and data services to become independent industries, but says, “The companies that understand data best are often the companies that understand AI best.”

20. Dexterous Hands Are a Joint Commitment to Data and the Robot Body

  • Chen Yilun is a “firm supporter” of dexterous hands. If the long-term end effector is a dexterous hand, the capture system should record the more than 20 degrees of freedom in a human hand, rather than forcing the human into a gripper or three-finger tool.
  • Gripper-to-gripper and three-finger-to-three-finger transfer are forms of dimensionality reduction. They make it easier to produce demos in the short term, but discard the full information required by real tasks. To truly break through a scenario, “not a single necessary link can be omitted.”
  • He cited Manus as a mature vertical player in motion capture and discussed Sandy Robotics’ Skill Capture Glove. Tashi Zhihang has also built a two-finger version, but Chen believes that if the team is serious, it should tackle all 5 fingers directly.
  • Responding to the criticism that Chinese teams lack landmark achievements, he argued that China had already produced end-to-end autonomous-driving results in 2021 and that embodied data capture is not lagging. Many domestic achievements simply have not received the same distribution.

21. China’s Advantage Comes From an Integrated Hardware-Software-and-Scenario System

  • Chen Yilun sees embodied AI as a repeated loop across hardware, scenarios, robot bodies, data, and algorithms. The modalities a model needs drive sensor design; data volume drives hardware cost; execution results feed back into the design of the robot body.
  • Many US startups struggle to seriously develop dexterous hands. Tesla is one of the few exceptions because it has an automotive industrial base. That is why Chen believes hardware capability cannot be separated from algorithmic competition.
  • His conclusion leaves little room for qualification: China and the US were roughly tied in the autonomous-driving era, but in the embodied era, “American entrepreneurs will not be competitors to Chinese entrepreneurs—not at all.”

22. AWE Puts World Evolution at the Center of the Model

  • AWE stands for AI World Engine. Tashi Zhihang allocates most of its neurons and compute to recording physical quantities such as time, space, and force, rather than the “retina-like representation” common in VLMs—local patterns, texture, and color.
  • This world representation must also capture interaction: after a robot squeezes an object, how does the object deform, and how does it push back on the robot? “Engine” emphasizes that the world evolves dynamically with each action and feeds back into the recommendation of the next action.
  • That is why Chen Yilun is not pursuing a mainline of “VLM plus an action head.” If a robot is merely a downstream head attached to a multimodal model, he argues, the industry is conceding that robotics does not deserve its own foundation model. His answer is the opposite.

23. A Robotics Foundation Model Should Be Rewritten From Robotics, Not Extrapolated From Image-and-Language Tasks

  • Chen Yilun’s method is to return to the classical robotics framework and redesign it with AI. The old theories may be dated, but they provide a scientific framework for perception, behavior, action, and control.
  • Language is valuable for describing high-level behavior and binding behavior to sensors and actions. Behavior is like semantics; low-level trajectories and control signals are like speech waveforms. Only by rising to the behavior layer does the system become capable of planning long-horizon tasks.
  • From his perspective, action is the lowest-level signal driving a physical system, although some in the industry define it as joint current and others as a trajectory. Behavior is an abstraction over a sequence of actions, analogous to strategy and tactics.
  • The model will remain closed-source in the short term to support efficient iteration. Chen’s standard for open source is not to “slap it onto the internet,” but first to let users access real value at low cost, and then allow more users to benefit from it.

24. 2026 Is About Data and Verticals; 2027 Is About End-to-End Payoff

  • Chen Yilun draws a parallel with the autonomous-driving cycle: teams seriously pursued end-to-end systems and data in 2019, then saw results in 2021. Embodied AI began in earnest in 2025, and his conservative estimate is that it “will definitely show results” in 2027; current progress may even be running ahead of that timeline.
  • He predicts a “data explosion” in 2026, followed by a surge in demo videos. More important will be the number of teams entering verticals honestly, because narrow scenarios increase data density and allow capabilities to be validated earlier.
  • Tashi Zhihang will not prioritize the consumer market. It will first target industrial manufacturing, where robots already exist and productivity needs are explicit—especially wiring-harness handling, connector insertion, and flexible assembly involving three-dimensional soft materials. These processes are widespread across autos, appliances, and servers, yet have remained difficult to automate.
  • Cheng Manqi noted that large-customer adoption is often tied to investment from industry players, making the depth of partnerships difficult for outsiders to judge. Chen’s answer is to look at the “value pool”: cooperation becomes substantive only when technology and products create gains that can be shared. To judge whether a company is credible, first ask whether it has truly figured out “what kind of self it wants to become.”