Pioneers Insight Method Research Author
134. A Review of Data with 谢晨: New Oil, Data Pyramid, and the Recipe
Back to Episodes

134. A Review of Data with 谢晨: New Oil, Data Pyramid, and the Recipe

Summary

  • 谢晨’s core judgment is that embodied data is not merely “in short supply” but structurally barren: if the data returned by 1M robots were a 60-point starting line, there may not be even 10K robots today capable of supplying enough real-world, simulated, or human data—“it might not even reach 0.6 points.” Real-robot data is the most accurate, but also the most expensive and hardest to scale across scenarios; the foundation that can push volume up the scaling law will be the two embodiment-agnostic categories: simulation and human first-person data. For investors, the value pool may therefore accrue not to the company with the most robots, but to an ecosystem spanning foundation-model companies, data engines, hardware vendors, and application owners.

  • The tightest bottleneck for the robot brain today is not collecting another batch of training material, but building evaluations that are difficult enough, repeatable, and scalable. Autonomous driving has shadow mode and LLMs have real user interactions, but robots cannot be tested every day across 1K homes and tens of thousands of tasks; BEHAVIOR’s current peak success rate is about 26% across 100 long-horizon tasks, showing that academic benchmarks have not been fully saturated, while industry needs much larger versions. “Without simulation, there is no way to run evaluation at scale.”

  • 谢晨 rejects “data is the model” as an end-state thesis: when a task distribution lacks a given category of data, models often genuinely fail, but long-term zero-shot generalization must be driven jointly by architecture and data scale, and generalization may only appear once data reaches sufficient mass. Over the past 6 months, foundation-model teams have started using robot arms as standardized testbeds, validating zero-shot transfer with embodiment-agnostic data and large-scale simulation; robotics companies care more about wheeled, legged, and dexterous platforms and delivery into specific hotels, factories, and solar settings. The two sides have split from the same starting point into two capital and product paths: “push generalization” and “deliver applications.”

  • The most counterintuitive recipe for high-quality data is not a perfect demonstration, but a correction trajectory in which the robot “fails first, then succeeds.” When making a pizza, a mushroom that is not gripped firmly, falls onto the table, and is then picked up and returned teaches recovery better than a flawless run; the distribution of bottle-grasping postures, partial successes and failures within long-horizon tasks, and precise evaluation matter just as much. Embodied data can cost anywhere from tens of yuan to several thousand yuan per hour, with high-quality data generally priced in the hundreds to more than 1K yuan; the spread reflects physical fidelity, trajectory expertise, and evaluation structure—not simply duration.

  • The embodied landscape will not simply replay Tesla: if more than 90% of autonomous-driving data comes from the company’s own fleet, the Tesla or Waymo model follows naturally, but robots face far greater complexity in tasks, objects, and physical interaction. 谢晨 therefore expects a Tesla-style integrated monopoly in robotics to be harder to build; the brain is more likely to come from model companies with LLMs, world models, VLA, RL infrastructure, and tens of thousands of GPUs, while embodiment companies handle manufacturing and deployment. “Without the simulation and human data at the bottom of the data pyramid, general embodied intelligence will not emerge.”

  • World models, VLA, and LLMs are forming a symbiotic stack rather than immediately replacing one another: world models focus on physical understanding and prediction, VLA on action at the edge, and LLMs on digital-world capabilities and the foundation layer. The same BEHAVIOR evaluation system is already being used to assess both VLA and world models, suggesting they could eventually converge; for now, world models need simulation for physical grounding, while simulation can use world models to expand generation and generalization. Tradable infrastructure therefore includes not only models and compute, but data systems that continuously close the loop across the physical world, action trajectories, tasks, and success conditions.

  • Big Tech’s investment in robot brains has clearly accelerated, with 谢晨 naming ByteDance, Alibaba, OpenAI, DeepMind, and Nvidia and grouping PI among the Frontier Lab competitors. China’s advantage remains concentrated in embodiment and manufacturing: Unitree’s positioning is the clearest, while AgiBot emphasizes integration across the supply chain; but base-model capabilities such as Qwen, talent density, and infrastructure mean Chinese brain teams “have a very good chance of catching up.” The end state is more likely to be collaboration among multiple powerful brains, data engines, embodiment companies, and application owners than a single model company taking everything.

Deep dive

1. 谢晨 Found the Industrial Prerequisite Through Repeated Trial and Error

  • 谢晨 studied physics at Peking University. His cohort started at about 110 students; he regularly worked until 2 a.m. for 3 years and stayed at school through winter and summer breaks, eventually finishing in the top 5. The experience taught him that effort can genuinely improve performance, while also forcing him to acknowledge that talent still sets the ceiling.
  • From a quantitative-finance PhD at Columbia Business School to dynamic pricing in e-commerce, product management, and autonomous-driving simulation, he repeatedly passed on directions that lacked innovation, social contribution, technical difficulty, or a commercial loop. The eventual point of differentiation was “building a product on a more disruptive technology, and using that product to genuinely support an industry.”
  • The search was not merely career planning. As an undergraduate, he organized an overseas-exchange program; during his PhD, he taught himself design and programming to help his pug, 土豆, which had heart disease, and built a dog-owner social app that ranked roughly in the top 3 in North America after downloading about 500 apps.
  • The dog-owner project ran for about 3 years and even received a term sheet from a Silicon Valley VC, but he shut it down because he had not worked out the business model and did not want to waste investors’ money. 张小珺 summarized his pattern as continually eliminating things that did not suit him. 谢晨 agreed: “I may simply not have been that lucky; it took me a long time to discover what I was good at.”

2. Cruise Turned Simulation from an Investor Demo into a Training Data Source

  • Before 谢晨 joined Cruise in 2018, its simulation system was more like “a toy” or a demo for investors: the game engine generated large volumes of data that looked realistic, but after the perception team used it, model performance fell rather than improved.
  • Cruise’s CEO Kyle gave him roughly 3 months to solve the problem. 谢晨 did not start by chasing more realistic visuals; he first established evaluation criteria for simulation, then combined generative AI, simulation, and algorithm iteration until synthetic data actually improved the model.
  • The validation changed his view: “Simulation and synthetic data are not a toy.” He initially called simulation a “time machine”—without it, autonomous driving might take 15 years; with it, the target might be reached in 5.
  • Real vehicles still supply most autonomous-driving data. Simulation mainly fills corner cases and provides repeatable regression tests, making it an accelerator in automotive rather than an irreplaceable prerequisite.

3. Nvidia and NIO Added Supplier and OEM Perspectives

  • 谢晨 left Cruise for Nvidia to challenge himself: the Waymo and Cruise approaches had not converged, and one company’s experience was insufficient to prove that his own method was best. A supplier’s vantage point allowed him to observe how multiple customers used simulation.
  • After joining Nvidia in 2021, he found that the largest customers for its in-vehicle Orin chip were not Waymo or Cruise but “NIO, XPeng, and Li Auto.” The realization led him to conclude that China might be the next center of autonomous driving, and he moved back with his family just 6 months after joining.
  • Inside Nvidia, he finally understood that the company was not a “gaming-card company,” or even merely a GPU company, but “an accelerated-computing platform company, a full-stack company.” He does not regret leaving, but retained that hard-tech judgment.
  • Back in China, he built a simulation-data loop at NIO from the OEM perspective, covering synthetic-data training, large-scale evaluation, and deployment. Holding positions at a supplier, an L4 company, and an OEM also helped him answer why he needed to work outside a single company.

4. Embodied Simulation Must Become External Infrastructure to Attract the Best Talent

  • 谢晨 co-founded Lightwheel AI with 杨海波 in 2023. The goal was not to build another autonomous-driving simulation vendor, but to become the data infrastructure and data engine for robotics.
  • His organizational view is that infrastructure is best built independently when the problem is sufficiently difficult and the market sufficiently large. The best algorithm talent at Cruise typically went into perception or prediction; the best data talent at Waymo did not necessarily join data infrastructure. An independent company can concentrate top algorithm, physics, and data talent on one problem.
  • Scale AI is his analogy: only when data production itself is complex enough can an external platform build cross-customer processes, feedback, and flywheels. If simulation serves only one autonomous-driving system, the case for an independent company is much weaker.

5. Data Is Evolving from a Static Textbook into a Model-Specific Education System

  • 谢晨 uses education as an analogy for data. In the ImageNet era, companies bought textbooks once: data consisted of static images, ground-truth labels, and fixed evaluation sets, resembling “cramming education.”
  • Scale AI industrialized the process through tooling, labor operations, quality control, efficiency, and delivery timelines, turning autonomous-driving labeling into a mass-market education factory.
  • In the LLM era, internet pretraining data is close to exhausted. The focus has shifted to post-training and evaluation; engineers, physicists, math-competition medalists, lawyers, and doctors act like advanced teachers who set questions, grade answers, identify weaknesses, and provide targeted experience.
  • His broader definition of data is therefore “signals that help you learn, together with the transmission of corresponding experience.” Data vendors are becoming teachers that understand the model’s state, proactively evaluate it, and stimulate new demand rather than merely taking orders.

6. Data Work Has Moved from Drawing Boxes to High-Priced Expert Feedback

  • The traditional autonomous-driving workflow covers sensor-data cleaning, slicing, box drawing, class and temporal labeling, followed by auto-labeling and human-in-the-loop quality control replacing pure manual work. 谢晨 estimates that 100K to several hundred thousand people may still participate in the industry globally.
  • New-generation suppliers such as Mercor and Surge provide models with questions, answers, comparisons, and RLHF feedback; participants are often highly trained professionals earning more than $100 per hour.
  • The key change is that they no longer simply append labels to existing data; they directly generate questions, demonstrations, and evaluations. For programming, the same task may have 10 valid implementations, and experts must identify which are good, poor, or ambiguous and explain the logic.
  • Unlike machine vision, which seeks a unique correct answer, LLM and embodied learning need to preserve human diversity, distributions, and even errors. “Strictly correct” and “strictly perfect” are no longer the only objectives.

7. “Fail First, Then Succeed” Teaches More Than a Perfect Pizza

  • Lightwheel’s initial request was for a robot to complete a long-horizon task perfectly in simulation: take a pizza base from the refrigerator, add toppings and cheese, put it in the oven, and press the button. Only a fully correct run counted as usable data.
  • Iteration with leading embodied-AI customers showed that correction trajectories were more effective: a mushroom is cut, not gripped firmly enough, falls onto the table, and is then picked up and returned to the pizza. These “negative examples,” or correction data, teach the model to recover rather than merely reproduce a successful path.
  • 谢晨 attributed the shift to stronger generalization: “The model is increasingly able to learn from mistakes.” 张小珺 said this is closer to how humans learn, and 谢晨 agreed directly.
  • The same applies to picking up a bottle: the broader the distribution of grasp directions, positions, and movements, the more valuable the data. The recipe is moving from one standard answer toward multiple paths, local failures, and explainable corrections.

8. “Data Is the Model” Describes the Present but Cannot Replace the Architecture Question

  • The view relayed by 张小珺 is that if a task category is absent from the model’s distribution, the task is hard to complete; capabilities emerge only after data has been “compressed” by the model, hence “data is the model, and the model is the application.”
  • 谢晨 acknowledged that this accurately describes current limitations in generalization: if an embodied model has never seen pizza-making, seeing vegetable cutting and hamburger-making may still not be enough to complete pizza-making zero-shot. At present, the desired task success rate often requires adding data for that task.
  • His objection is long term: a model whose architecture lacks zero-shot transfer is not a real path to general intelligence. Data may need to reach sufficient scale before generalization appears, but architecture still determines whether the capability can emerge.
  • He compares it with human learning: an ordinary person and Elon Musk may receive similar information, but the latter is better at transferring knowledge quickly through first-principles reasoning, broad knowledge, and practice. The model’s own generalization ability still needs to improve.

9. Over the Past 6 Months, Foundation-Model Teams Have Made Zero-Shot the Primary Goal

  • 谢晨 observed that about 6 months ago, foundation-model and robotics customers were still fairly similar in how they thought about data volume and definition. Over the most recent 6 months, a “qualitative change” emerged, with the former concentrating on validating zero-shot generalization.
  • These teams believe that sufficiently effective algorithms, enough high-quality embodiment-agnostic simulation and human data, and large-scale simulation evaluation can let a standard robot arm transfer to unseen tasks.
  • They deliberately use robot arms and grippers rather than wheeled, legged, or humanoid platforms—not because they want to build hardware, but to reduce maintenance, debugging, and embodiment differences and isolate brain capability as the main variable.
  • The goal is not to tighten a particular screw immediately, but to ask whether a model trained on 10 or 100 tasks can complete another 5 it has never seen. “Generalization” itself is the product.

10. Robotics Companies Are Moving in the Opposite Direction, Deeper into Embodiment and Specific Scenarios

  • Unlike foundation-model teams, robotics customers are increasingly focused on the complexity of wheeled and legged platforms, hand sensors, and defined businesses such as hotels, auto workshops, supermarkets, and solar-panel maintenance in deserts.
  • They need to know whether a specific task can be executed reliably and whether the economics work, rather than first proving abstract zero-shot capability. The two customer groups have therefore split from similar data buyers into “pushing the scaling law” and “delivering scenarios.”
  • Foundation-model teams prefer readily available data from homes, supermarkets, and some factories to expand general-purpose understanding; robotics companies customize data around their first path to commercialization.
  • 谢晨’s language remains stage-specific: he believes embodied zero-shot is “gradually beginning to emerge,” which makes him optimistic, but he does not claim that the generalization problem has been solved.

11. LLMs, World Models, and VLA Form a Three-Layer Collaboration Stack within One Company

  • Most companies put LLMs, world models, and VLA in separate teams, but 谢晨 considers them “extremely symbiotic”: VLA can reuse the company’s own foundation model, world models fill the gap in physical prediction, and VLA feeds action outcomes back into the world model.
  • Robotics teams without a strong foundation can only plug into open models such as Qwen. Companies with globally leading foundation models can share data understanding, correction experience, and training infrastructure across the stack.
  • The resource gap is especially visible in GPUs and RL infrastructure. A robotics company having several thousand GPUs is already substantial; a foundation-model team may have tens of thousands, a gap of at least an order of magnitude.
  • Large-scale parallel RL infrastructure is difficult to build from scratch for a single embodied project. Foundation-model companies can transfer existing systems to VLA fine-tuning, creating a practical barrier to entry into robot brains.

12. World Models Focus on Prediction, VLA on Action, and Their Evaluation Standards Are Converging

  • 谢晨 sees the world model as a possible cloud-based brain that understands and predicts the physical world; VLA is closer to an edge brain, requiring precise, effective, and efficient action; LLMs already possess some world-model capabilities in the digital world.
  • Adding an action head to the same base could bring a world model into VLA tasks. Conversely, VLA deployment feedback can test whether the world model truly understands physics.
  • He points to BEHAVIOR as the connection point. The simulation benchmark advanced by 李飞飞’s team contains difficult long-horizon tasks, and its leaderboard includes not only VLA teams but also teams starting with world models plus action heads.
  • Another company, Inact, evaluates world models against similar standards. 谢晨’s inference is that “if the evaluation systems become increasingly consistent,” the two may eventually merge, although they remain mutually dependent in the near term.

13. The Embodied Landscape Will Be Built by Four Parties: Brains, Bodies, Data, and Applications

  • The premise of Tesla’s data engine was that Tesla may already have had more than 1M vehicles continuously returning driver data. The cloud brain improved and was then deployed to the edge, allowing the embodiment company to control both the largest data pool and the best model.
  • Robots have neither millions of devices operating all day nor low-cost, scalable human teleoperation, so most data will not be generated by the embodiment company. 谢晨 therefore believes the Tesla loop is “unlikely to hold at the foundation” in embodied AI.
  • Foundation-model companies will use embodiment-agnostic data to push generalization, then hand the brain to embodiment companies for fine-tuning and deployment. Data companies will provide evaluation, feedback, and experience, while application owners control real demand in factories, healthcare, agriculture, and other settings.
  • Application owners can choose between hardware A and B, or even develop robots themselves like OEMs, giving them substantial bargaining power. Their understanding of mass production, quality, stability, and cost makes them more than end customers: they are the ecosystem’s fourth party.

14. “System-Level Capability” Is Closer to the Essence of a Model Than Knowledge or Data

  • 张小珺 asked: if one cannot say “data is the model,” then what is a model? 谢晨’s answer was not “knowledge is the model,” but continuously improving system capability.
  • A child reads picture books; a mature expert receives more targeted, higher-order signals. Every upgrade in learning ability generates new data requirements. A stronger model therefore will not simply need less data; it will change the level and form of the data it needs.
  • He accepts the “private tutor” analogy, but stresses that the tutor must be organized around a system rather than an unlimited accumulation of people. Only then can “teaching by word and example” scale.

15. LLM Data Is Around 60 Points; Embodied Data on the Same Scale May Be Below 0.6

  • Internet pretraining data for LLMs is close to “topped out.” The problem has shifted mainly to post-training and evaluation: finding more advanced people, setting harder questions, and providing finer-grained feedback.
  • 谢晨 reluctantly estimates the current LLM data state at roughly 60 points on a 100-point scale, while emphasizing that post-training and evaluation still have substantial room to improve. He also notes that human learning is endless, making “100 points” difficult to define.
  • For embodied AI, he treats the data generated by 1M robots as only a 60-point starting point. The combined equivalent of real, simulated, and human data may not even reach 10K robots, hence “below 0.6 points.”
  • The gap is not only about volume. Embodied-data collection must simultaneously handle physical scenes, interactive assets, action experience, language definitions, and evaluation signals; the difficulty may be several orders of magnitude higher than for LLMs.

16. Autonomous Driving and LLMs Both Have Nearly Free Shadow Modes; Robots Do Not

  • Autonomous-driving teams can deploy a new algorithm in shadow mode on the vehicle without actually controlling it, then compare the model’s output with the driver’s actions. The difference provides both a cheap evaluation signal and a human demonstration.
  • Once an LLM is deployed, its interaction with users is itself a form of free feedback. Users demonstrate and provide feedback in different ways as they use the model, helping it improve.
  • Robots have not yet been deployed at scale, and the real world does not allow immature policies to trial and error in parallel safely and cheaply. The equivalent evaluation loop is therefore missing.
  • 谢晨 sees simulation as the only scalable substitute: let the brain undergo repeated tests in parallel physical worlds, receive success and failure signals, and send the feedback back into training.

17. Digital Agents and Physical Robots Face the Same Learning Structure

  • 张小珺 noted that LLMs also have not seen real human work after moving from chatbots into agents, bringing the data problem closer to robotics. 谢晨 summarized it succinctly: “A robot is an agent in the physical world.”
  • Both require an environment, transmitted experience, and evaluation. Digital agents use RL environments—virtual ride-hailing, e-commerce, shopping, or programming environments—to trial and error against success conditions and fine-tune repeatedly.
  • Physical agents will ultimately also use simulation environments to run RL around defined tasks, scenes, and success metrics. The difference is that one environment simulates web and software interactions while the other simulates physical interaction.
  • But 谢晨 is clear on priorities: embodied AI still lacks a pretrained foundation and scalable evaluation. RL post-training has begun, but for now it is the “second-order problem.”

18. BEHAVIOR’s 26% Success Rate Shows That Evaluation Still Has a Real Slope

  • Other academic-level embodied benchmarks have largely been “blown out” by the leading foundation-model customers 谢晨 serves and can no longer distinguish models. BEHAVIOR may contain 100 long-horizon tasks, with 谢晨 citing a current peak success rate of about 26%.
  • The benchmark matters not only because the score is low, but because it provides difficult, reproducible tasks that are hard to collect at scale on real robots, allowing model capabilities to be compared continuously.
  • Academic BEHAVIOR is still not the same as industry evaluation. Foundation-model teams need thousands of scenes, 1K to 10K tasks, and continuously updated success conditions to evaluate generalization in open environments such as homes.

19. The Data Industry Has Shifted Three Times with the Learning Paradigm: Textbook, Factory, Feedback Engine

  • 李飞飞 and ImageNet established the static form of a training set plus evaluation set. Scale AI, riding the rise of autonomous driving, turned labeling into an industrial production line that could control quality, efficiency, and timelines.
  • 谢晨 views Scale’s data factory as something like a wafer fab: the real secret sauce is process, standards, know-how, and execution—not a simple labor pool.
  • He defines the next phase as evaluation-driven. Data vendors first help models identify problems, then stimulate new demand, deliver targeted experience, and iterate the production chain using model feedback.
  • 谢晨 said Scale entered what he calls the “GPT-2 and RLHF” data phase around 2021–2022; the customer relationship shifted from buyer and vendor to partnership.

20. The Number of Experts Has Not Fallen as Models Improve; Demand May Continue Rising

  • Although hourly pay for individual data contributors has risen sharply, the number of industry participants has not declined. 谢晨 initially expected better learning efficiency to reduce the need for experts, but reality has not supported that view so far.
  • He compares it with test-time scaling. When DeepSeek emerged, some assumed pretraining and demand for Nvidia GPUs would decline; instead, more agents and applications stimulated more compute consumption.
  • Similarly, “the more capable a person is, the more they like to learn.” As models improve, they may require more and higher-order knowledge rather than less data. If AI ultimately reaches Nobel Prize level, almost no one may be able to teach it directly.
  • At that point, learning would shift from external teachers to self-comparison: the system would need an environment, continuously updated success criteria, and RL to become “better today than yesterday.”

21. For Robots, Simulation Is a Requirement, Not an Accelerator

  • 谢晨 gives a direct conclusion: “Without simulation, this simply cannot be done.” Robots cannot collect enough data at scale through real embodiments, and large-scale evaluation has no viable source other than simulation.
  • 10 or 20 robots can support laboratory testing, but not repeated daily testing across 1K homes and thousands of tasks with precise comparisons between every algorithm version.
  • Human first-person data can supplement training demonstrations, but cannot replace controlled, reproducible evaluation at scale. Simulation is therefore the middle layer of the training pyramid and almost the only scalable infrastructure for evaluation.

22. Frontier Lab, Once Committed to Real Robots, Turned to Simulation Because of the Evaluation Bottleneck

  • Lightwheel’s early customers were mostly committed simulation believers. Other top teams explicitly said they knew Lightwheel’s simulation capabilities were strong, but “the time has not come yet.”
  • Over roughly the past 3 months, those teams began approaching Lightwheel proactively, primarily because they could not evaluate at scale.
  • A household-robot team may already know how to fold clothes and perform chores. The next step is testing across 1K homes, continuously changing tasks, and evolving evaluation standards—something real robots cannot provide.
  • Robotics companies have long used small-scale RL simulation for locomotion and full-body control, often running it on one local machine. The new requirement is industrial-scale data and evaluation.

23. China’s Real-Robot Camp Reflects Both Technical Judgment and Business-Model Constraints

  • 张小珺 observed that Chinese robotics teams appear more likely to favor real robots than simulation, often claiming that real-robot data generalizes more easily. 谢晨’s first explanation was: “Your position determines your view.”
  • Many companies still make money by selling embodiments or “data-collection centers”: customers buy robots to collect data. If a company publicly argues that simulation is sufficiently effective, it becomes harder to persuade customers to keep buying large fleets for collection.
  • He does not deny the need for real-robot data and believes multiplying the current scale by 10x could still be reasonable. The debate is over what share it ultimately occupies in the data pyramid, not whether real robots are needed.
  • Real-robot data centers themselves often amount to “simulation of the real world”: fixed tables, IKEA-style scenes, and fake bananas and apples. The problem is that scenes change slowly, making it hard to cover the breadth and diversity of the real world.

24. Sim-to-Real Is a Domain Gap, Not Generalization Itself

  • 张小珺 cited 谭杰’s distinction: simulation data creates a Sim-to-Real problem, while generalization should be solved through sufficiently large volumes of simulation data. 谢晨 explicitly agreed.
  • If large amounts of simulation and real data are mixed during pretraining, the model’s general capability may improve and the simulation–real gap may gradually narrow. The current gap cannot itself prove that simulation lacks generalization.
  • 谢晨 believes many robotics companies have not actually trained pretraining-scale foundation models. If the goal is only a specific embodiment and task, naturally less data is needed; that does not disprove embodiment-agnostic scaling laws.

25. Reproducibility, Intervention, and Physical Accuracy Define the Boundary of Simulation

  • 谢晨’s definition of simulation is demanding: actions must be executed in a sufficiently physically accurate environment in a reproducible and correctable way, with their outcomes observable.
  • “Physical accuracy” is not limited to geometric and visual similarity; it includes parameters such as friction. “Reproducible” does not mean exactly identical 100 out of 100 times, but the same conditions should produce the same result 95 or 99 out of 100 times.
  • Intervention is even more important. Holding the environment and initial state constant while changing only the action should reliably produce a different outcome. Only then can data support reliable comparison of action outcomes and large-scale evaluation.
  • Ordinary video-generation models mostly predict the next frame and often lack reproducibility and sufficiently accurate action control. They therefore cannot all be broadly labeled simulation today.

26. World Models Will Not Eliminate Simulation; the Two Will More Likely Reinforce Each Other

  • World models have stronger generation capabilities and can cover broad, relatively realistic world predictions, potentially incorporating robot actions over time. 谢晨 therefore believes they may eventually become “a type of simulation.”
  • But world models still need simulation to improve physical grounding through accurate physics, real interaction, and human-like actions. Otherwise, generated outputs may not be controllable or evaluable.
  • In the other direction, world models can help simulation expand scenes, improve generation, and strengthen generalization. The relationship between Lightwheel and its world-model customers is: “They use our data, and we use their models.”
  • 张小珺 tried to classify simulation as a method within world models, but 谢晨 did not accept the framing: neither is a subset of the other; their shared objective is to provide better learning capabilities for intelligence.

27. Robotics Will Not Simply Recreate the Waymo–Tesla Contest

  • Autonomous driving has relatively simple tasks: the vehicle–road physical relationship is limited, and when confronted with a cup, the car mainly needs to avoid it. A robot must judge material, size, and grip force before selecting an action.
  • Autonomous driving may compress lower-ceiling intelligence into the vehicle architecture through large volumes of embodiment-specific data and imitation learning. Removing language would, in 谢晨’s view, “dramatically reduce” intelligence, but driving itself might still work.
  • Another path is a more general VLA that learns to drive alongside other tasks. 谢晨 does not declare either path the winner, leaving open the possibility that both can work.
  • Because embodied data is primarily generated at the embodiment-agnostic layer, Tesla’s or Waymo’s organizational model cannot be applied directly. Even if Optimus pursues integration, the body and brain may still reside separately at Tesla and xAI.

28. Vertical Robotics Can Succeed, but It Looks More Like Waymo Than a General-Purpose Brain

  • 张小珺 proposed that robots could, like autonomous driving, collect large amounts of real-robot data in a single vertical and train the scenario first. 谢晨 acknowledged the path and compared it with Waymo and Cruise.
  • He experienced this firsthand at Cruise’s initial deployment in San Francisco. Even the first city was extremely difficult; moving to a second city still required new data collection, training, and large-scale evaluation. This is not a strong-generalization route.
  • A vertical robot company can deepen one or two scenarios, build a business model, and create a moat. He considers mining autonomy a valid example. But transferring to another scenario may be “bone-breaking,” requiring changes to both model architecture and data.
  • He therefore believes the market underestimates both the difficulty of vertical deployment and the difficulty of cross-domain transfer after deployment. Generalization in a general-purpose model may emerge earlier instead.

29. The Data Pyramid Places Real Robots at the Top, Not at the Broadest Base

  • The concept, proposed by Professor 朱义可, a student of 李飞飞, has 3 layers: real-robot teleoperation data at the top, simulation data in the middle, and internet and human video at the bottom.
  • Real-robot data is the most accurate and useful, but hard to expand across robot counts and scenes. Simulation scales but must address Sim-to-Real; human first-person data is the largest and embodiment-agnostic, but requires reconstructing actions and physics.
  • Evidence cited by 谢晨 includes the BEHAVIOR Challenge, Nvidia’s GR00T, which uses large volumes of simulation data, and Generalist, which uses about 270K hours of data collected through human operation of five-finger grippers.
  • He believes these results already show early signs of an embodied-data scaling law. Over the past several months, Lightwheel has also shifted from “stimulating customer demand” to requiring a scaled team capable of actually delivering against it.

30. Each Layer of the Pyramid Must Still Be Split by Quality and Scale

  • At the top of the simulation layer, humans can teleoperate robots. The advantage is high-quality demonstration without dependence on real robots; the drawback is that scale remains constrained by human labor.
  • Further down, models or algorithms drive automated collection. Human intervention is limited, and scale rises substantially, but quality is generally lower than direct human demonstration.
  • Human data also divides into active and passive data. Large-scale passive data from people casually wearing glasses offers volume but weak quality control; specialized hardware and processes improve quality at the cost of scale.
  • The data pyramid is therefore not a ranking of which category is absolutely best. It is a portfolio allocation across realism, action quality, coverage, cost, and scalability.

31. The Real Data Loop May Rotate Around Simulation and Evaluation

  • 谢晨 corrected the impression of a static pyramid: the 3 layers are not independent, but form an evaluation-driven loop centered on simulation.
  • Building credible simulation evaluation requires real-world scenes, physics, human trajectories, tasks, and success criteria. A simulation team working in isolation cannot define what counts as success in the real world.
  • Real-to-sim reconstructs scenes, objects, tasks, and standards from video into a parallel world. Simulation results then return to real robots for comparison with real teleoperation and real-world evaluation.
  • Sim-to-Real therefore serves both training and evaluation: the objective is not only to make a policy deployable, but also to prove that simulation scores correlate with real performance.

32. Human First-Person Data Treats People as the Largest Robot Fleet

  • When foundation models pursue cross-embodiment capability, humans themselves can be viewed as a type of robot. Adding human first-person actions to training is equivalent to “treating people as vehicles.”
  • The closer the camera is to the eyes, the better. Chest-mounted devices and head-top cameras diverge from the human visual-action relationship, making smart glasses a more natural collection terminal.
  • 谢晨 believes human-data companies should not build niche hardware themselves because they cannot achieve consumer scale. Consumer-grade products can reach 1M or more units; the real breakthrough is to use devices consumers already want to wear and design the collection workflow through an SDK, API, or app.
  • Meta Ray-Ban is the product concept he endorses: first build an everyday pair of glasses that looks good enough, then add a camera and AI assistant. The ideal is not to wear glasses for the sake of data, but to like the glasses and contribute data incidentally.

33. Real-Robot Data Is Overvalued; Simulation Evaluation and Human Data Remain Undervalued

  • 谢晨’s first judgment is that real-robot data is “definitely overvalued.” Industry developments have partly confirmed it: teams once committed to real robots are beginning to buy simulation data, simulation evaluation, and human data.
  • The value of simulation for training is gaining recognition, but the value of scalable evaluation remains underappreciated. Foundation-model teams already feel the pain; many robotics companies will encounter it only after their task, scene, and model-version counts expand.
  • Human data is similarly undervalued, although its standalone moat is limited. In 谢晨’s framework, its value is to add the real world, human experience, and evaluation standards to a simulation-centered loop.

34. Embodied-Data Pricing Depends on Learning Value, Not Production Cost Alone

  • 谢晨 believes data is becoming “more expensive overall” because the marginal value delivered to algorithms differs completely between static labels and high-order feedback.
  • Pretraining data is closer to a commodity and should have its cost shared among a handful of global model companies because it improves general foundation capabilities. Post-training and evaluation are highly model-specific, making their value and price higher.
  • His market range is broad: structured embodied data can cost from tens of yuan to several thousand yuan per hour, with high-quality data generally priced from several hundred to more than 1K yuan.
  • A data point contains at least 3 elements: the physical scene, action and language experience, and evaluation of success or failure. Long-horizon tasks must also label the success or failure of local steps, rather than providing only an end-state label.

35. High-Quality Data Needs Three Kinds of Fidelity: Physics, Trajectory, and Evaluation

  • Scenes must be sufficiently diverse, and object interactions must follow real physics. Data that merely looks realistic while getting friction, flexibility, or contact wrong cannot support reliable training.
  • Trajectories must be professional and natural, including smooth operation, reasonable mistakes, and subsequent correction. In pizza-making, a clip showing food being dropped and then retrieved may cost more than a perfect video.
  • Evaluation and semantic labels must be accurate. Long-horizon tasks in particular require fine-grained state decomposition; human video also requires sufficiently faithful hand and full-body tracking.
  • Movies and ordinary video are not useless, but their ROI is lower: 2D footage requires additional labeling, consumes substantial compute, and may deliver limited gains in intelligence.

36. Game Data Is Better Suited to World Models, but Not Necessarily the Highest-Value Layer

  • Games provide 3D environments and large volumes of executable trajectories. World-model teams may even buy game rights, have agents play automatically, and collect the resulting data.
  • The problem is cross-domain transfer: the physics, assets, and tasks in a game may diverge from reality. Usefulness does not mean maximum utility.
  • 谢晨 puts the higher-ROI pretraining mix today in algorithm-driven simulation data with humans in the loop, plus high-quality human first-person data. The data pyramid is broad, and a supplier does not need to cover every raw material.

37. Lightwheel Prefers “Data Engine” to “Data Factory”

  • To 谢晨, “data factory” remains a production-line metaphor: standardized output, weak feedback, and limited room for technical or systems-level improvement. A data engine discovers problems through evaluation, then uses the feedback to drive the next data cycle.
  • Lightwheel has about 100 full-time employees, mostly engineers and technical staff. Its core is not a single labeling team, but a full-stack system spanning physics simulation, automated measurement, algorithmic collection, semantic labeling, and real-world evaluation.
  • Rigid-body assets are relatively easy. Non-rigid tasks such as inserting and removing cables require an in-house physics solver and joint asset calibration; the team also uses robot arms to automatically measure the mechanical properties of real objects and write them back into simulation assets.
  • 谢晨 rejects the perpetual-motion idea that “AI generates all its own data.” The system still needs an accurate world, real tasks, and human experience; technology’s role is to amplify those external signals.

38. From Human Teleoperation to Automated Collection, the System Could Amplify Experience by About 2 Orders of Magnitude

  • One Lightwheel pipeline has humans teleoperate simulated robots. It can use different embodiments or standardized custom bodies to collect high-quality demonstrations.
  • On top of that, automated algorithms collect trajectories at scale, with human intervention limited to a few stages. Foundation models then perform semantic labeling, while human-in-the-loop quality control protects output quality.
  • Another evaluation pipeline starts with human video and real tasks, reconstructs physics, extracts tasks and success criteria, and places them into a scalable simulated world.
  • If everything were centered on humans, 谢晨 estimates that 10M to 100M participants might ultimately be needed. The amplification effect of systems and simulation could reduce the labor requirement by roughly 100x.

39. Evaluation Engineering Must Be Difficult Enough Yet Replicable at Scale

  • A clothes-folding demo in one fixed scene cannot test generalization. Real evaluation may require thousands of parallel scenes, 1K to 10K tasks, and explicit success conditions.
  • Simulation and real-to-sim can expand scenes and physics. The harder problem is extracting tasks and evaluation standards from the real world; if evaluation is disconnected from reality, greater scale only makes it less meaningful.
  • Lightwheel maintains real robots and real-world evaluation infrastructure not to perform a small number of physical tests for customers, but to verify whether the same algorithm’s real and simulated performance correlate.
  • Only with that correlation can simulation evaluation avoid becoming a system that “sets its own questions and gives itself high scores” and instead serve as a credible dashboard for customers to compare model versions every day.

40. Data and Model Companies Must Find the Recipe Together Rather Than Shift the Blame

  • The finger-pointing described by 谭杰 is real: data vendors say the model was not trained properly, while model companies say the data is poor. 谢晨 compares the current phase with Scale AI and OpenAI jointly discovering the data recipe in their early days.
  • The broad direction is becoming clearer—simulation, human data, and evaluation—but the details are still changing: from perfect trajectories to correction data, and from one grasping style to a broader action distribution.
  • Only about 5 teams in the world may be able to validate a pretraining-scale embodied-data recipe. Lightwheel works with several of them and is continuously validating the data pyramid with roughly 2 teams.
  • The experiments require tens of thousands of GPUs to compare the mix across layers, the handoff between pretraining and RL post-training, and how simulation, real data, and evaluation fit together into one integrated recipe.

41. The Secret Sauce of Data Companies Is Increasingly Like Pedagogy

  • 谢晨’s most counterintuitive realization is that data no longer has one unique correct answer. Effective learning is closer to letting the model see errors, understand corrections, and reach its own conclusions from multiple solutions.
  • Having one teacher demonstrate every question may not be optimal. Treating every student as a teacher and observing different approaches to the same question may create stronger generalization. Distribution and causal structure matter more than a single “perfect answer.”
  • He therefore says: “At the end state, data companies may look a lot like education companies.” Embodied AI still needs substantial imitation and demonstration today; over time, it will need more challenges, physical interaction, and targeted evaluation.

42. Big-Tech VLA and World-Model Teams Are Converging on Embodiment-Agnostic Approaches

  • 谢晨 calls one group the “foundation-model camp”: Big Tech VLA and world-model teams care most about zero-shot capability, are less concerned with making the embodiment complex up front, and use standardized robot arms to validate upper-layer capability.
  • These teams believe in simulation, human data, and simulation evaluation, and are also moving earlier into large-scale RL. The logic is inherited directly from the LLM scaling law.
  • 张小珺 asked whether Big Tech would still prioritize LLMs. 谢晨 acknowledged that this was more realistic 3 to 6 months ago, and even before this year, but said teams have begun freeing resources for robotics as the language-model trajectory has become relatively clear.
  • He later adjusted the window to “close to the past year” and emphasized that the inflection point came as the industry increasingly recognized that if the core data is embodiment-agnostic, the robot brain naturally becomes an opportunity for foundation-model companies.

43. Five Big Tech Companies and PI Are Accelerating the Race for Robot Brains

  • Asked who had become more aggressive, 谢晨 named ByteDance, Alibaba, OpenAI, DeepMind, and Nvidia in that order, saying they were “definitely more aggressive.”
  • He also grouped PI with them. Although it is a startup, its positioning is closer to a Frontier Lab than to a robotics company selling embodiments.
  • These players share the same assets: strong foundations, world models, VLA, RL infrastructure, and tens of thousands of GPUs. Their common objective is to first prove general-purpose brain generalization.
  • xAI also has a chance, but 谢晨 believes it has not yet won the LLM race, while Tesla must first realize the hardware advantage of Optimus. The two sides have not fully converged.

44. Embodiment Companies’ Advantage Is Manufacturing, Not Building a Unified Brain

  • 谢晨 likes Unitree because its positioning is clear: focus relentlessly on hardware that is stable enough for mass production, without competing with brain companies for the same layer of value.
  • When Big Tech brains need to enter real-world scenarios, Unitree could become a high-priority embodiment partner. “Knowing your boundary” is itself a competitive advantage.
  • He believes AgiBot has emphasized integration across the upstream and downstream chain since day one, with solid progress in commercialization and mass production. Embodied AI may still be a supply-driven market, where producing enough units first helps activate the supply chain and industrial base.
  • Figure wants to become the Tesla of embodied AI, covering hardware, mass production, deployment, and the brain. 谢晨 believes the challenge is extremely difficult and the scenarios remain unclear: “It is still far away.”

45. The End State Is More Likely Ecosystem Collaboration Than a Single Robot-Brain Monopoly

  • If the largest data pool were tightly tied to one embodiment, the company with the most deployments could build a Tesla-style monopoly. But because embodied data is primarily embodiment-agnostic, that premise is weakened.
  • 谢晨 therefore expects the best brain company, data company, and robotics-hardware company to collaborate closely, serving application owners that control real-world demand.
  • Some application owners may also become strong hardware companies, but a foundation-model company will still struggle to build a standalone monopoly. Brain competition may look more like today’s foundation-model market: many once assumed OpenAI would take everything, but it did not.
  • For startups, he believes training a unified brain from scratch is “not very rational.” Vertical tasks, hardware, data infrastructure, or application deployment are more consistent with their resource constraints.

46. China Could Close the Brain Gap, but the Route Has Not Fully Converged

  • The US currently leans toward robot brains, while China leans toward embodiments and mass production. 谢晨 still believes China “has a very good chance of catching up,” citing base-model capabilities such as Qwen, talent density, infrastructure, and Big Tech’s willingness to invest.
  • His potential leaders include OpenAI, DeepMind, Nvidia, and China’s ByteDance and Qwen teams. OpenAI’s existing robotics team “should not be underestimated.”
  • The path remains unsettled. Embodiment-agnostic data has shown early signs of a scaling law, but brain architecture, the incorporation of world models into VLA, and further gains in generalization remain open research questions.
  • He separates several concepts often conflated: Physical AI refers to autonomous-driving and embodied models capable of acting in the physical world; spatial intelligence focuses on 3D reconstruction, generation, and prediction; world models focus on physical understanding and prediction but may not possess action capability.

47. The Single Most Valuable Problem to Solve Is Scalable Evaluation

  • If he could solve only one embodied-data problem, 谢晨 would choose evaluation rather than further expanding ordinary data collection. The path and early scaling-law signals for embodiment-agnostic pretraining already exist; the bottleneck is still the dashboard that measures intelligence gains.
  • The solution must be a physically credible, sufficiently difficult, repeatable, expandable simulation evaluation that is continuously updated from real-world tasks and standards rather than a one-off benchmark.
  • For LLMs, he also places the bottleneck in evaluation and post-training: the stronger the model, the more capable the people, more difficult the questions, and more effective the evaluation metrics it requires.

48. Data Will Not Disappear; It Will Become the Environment in Which Models Train Themselves

  • 谢晨 once believed data might no longer matter 15 or 20 years from now, but has changed his view: the stronger intelligence becomes, the greater its hunger for knowledge and feedback may be.
  • The object of learning is changing—from learning from other people to comparing oneself with yesterday; from reading standard textbooks to practicing, failing, and receiving feedback in real or virtual environments.
  • The mass-market, standardized knowledge production represented by the data factory may become obsolete relatively quickly, but RL environments, simulated worlds, success conditions, and evaluation systems will remain first-order requirements.
  • He imagines an end state that may resemble what Elon Musk has said: “Maybe we humans are living inside a simulation.” AI repeatedly trains in an environment with defined metrics, while data suppliers evolve into environment and feedback-system providers.

49. Einstein’s Thought Experiments Explain Why Simulation May Persist

  • 张小珺 asked whether sufficiently powerful AI would still need an education system. 谢晨’s answer moved from “teachers” to “environments”: whether digital or physical, self-improvement still requires a context, constraints, and a definition of success.
  • Einstein did not rely only on external textbooks; he built thought experiments on existing physical knowledge and theorems. 谢晨 sees the process of setting premises in the mind, reasoning through them repeatedly, and trialing and erroring as a form of simulation.
  • Advanced intelligence does not need an infinite number of standard answers so much as enough environments, physical knowledge, and constraints to generate experiments and test conclusions on its own.
  • Returning to his own career choice, 谢晨 confirmed that simulation is the direction he had been searching for: it is the foundation and prerequisite of embodied learning, but not the only answer. The end state remains “a data pyramid centered on simulation, but not simulation alone.”