张巍 of LimX Dynamics on Embodied AI, World Models and the Robot Brain
张巍 of LimX Dynamics on Embodied AI, World Models and the Robot Brain
Summary
- 张巍’s core view is that embodied AI is “not competitive enough” yet: it is not a vertical market that can accommodate only one number one, but more like the entire internet, with an application breadth that could even far exceed that of EVs. Cars mainly serve the need to get from A to B, while embodied AI is a general-purpose vehicle that can penetrate every industry; its functions, objectives and users have yet to converge, and the sector has not entered a homogenized battle over price and efficiency.
- LimX Dynamics has explicitly taken bipedal humanoids out of the industrial-efficiency-tool category: industrial embodied AI will thrive, but “humanoids will not enter factories,” because factories are designed for machines while humanoids are designed for human environments. The product function behind “Serve people, not process” is “maximize number of tasks over form factor”; a single humanoid maximizes task coverage, while robotic arms, quadrupeds and modular purpose-built machines maximize efficiency within defined tasks.
- World models currently look more like a new hope for VLA data scaling than a major technological breakthrough that has already happened. They predict future states and observations based on the current state and action; in embodied AI, the focus is still primarily video, with the field now extending toward World Action Models that predict actions as well. Time-series video carries physical laws better than static frames, while ego-centric and existing internet video are easier to scale than real-robot data.
- Physical equations can in principle add information to world models, but the real bottleneck is alignment, not whether Newton’s laws are present. 张巍 sees physical equations as a form of data compression: “Newton’s laws are also a compression and representation of all motion data”; the practical route is often to use physical laws for simulation, generate visually aligned data, and then confront the difficult sim-to-real problem.
- Embodied models cannot copy the foundation-model playbook of going “general first, applications later”; general capability should emerge from a “flywheel between general models and scenario data.” Driving data and egg-peeling data may not belong in the same training run, and blindly stacking cross-skill data in the hope of emergence is “carving a boat to mark where the sword fell”; the near-term unit of deployment is an individual skill, and the key metric is whether the commercial value created by that skill can cover the cost of pretraining, post-training and scenario data.
- “A model is not a brain; the brain is an operating system”: LimX Dynamics breaks the robot brain into three layers— a cerebellar foundation model, Human Brain VLA and Agentic OS. The cerebellum executes whatever movement is requested, VLA acts as System 1 to turn vision, environment and action into higher-order skills, and CoSA acts as System 2, calling LLM, VLM, VLA, GPS and other models and tools for memory, reasoning, decision-making and task orchestration.
- Humanoid commercialization will not start with box moving, but proceed from performance, guidance and sales assistance toward physical manipulation. 张巍 divides applications into no interaction, weak interaction and strong interaction: performances already have sell-out demand; the next step is a mobile advisor that “talks but does not touch,” with the industry only later seeking strong-interaction skills whose data costs can be matched by commercial value; a unified body that continuously adds apps could eventually close the loop and enter homes filled with diverse, low-frequency tasks.
- Financing should serve a clear objective, and valuation should maintain a healthy relationship with value; going public does not necessarily mark the end of a boom. 张巍 believes embodied AI has “a higher ceiling than foundation models, and a higher floor,” and says academic founders must transform from academia into technology, engineering, product and commercialization; organizationally, the match between talent and roles changes by stage. 李翔 closed by summarizing his most important shift in perspective: “startups can die” and “if you cannot afford to lose, you cannot win.”
Deep dive
1. Embodied AI Is Far from a One-Winner Market
- Looking back over the past 2 years, 张巍’s takeaway is not that one route has won, but that the industry has been “constantly iterating and changing”; every time the market assumes the heat has peaked, more financing, attention and change follow.
- He rejects applying the number-one logic of an internet vertical market to embodied AI: it is closer to the entire internet, capable of accommodating “Meituan, Alibaba, ByteDance and Tencent” while producing different companies across a wide range of niches.
- The host compared this with the coexistence of multiple routes in EVs, but 张巍 believes the opportunity in embodied AI is “much, much larger than the EV sector”: cars mainly solve mobility, while robots are general-purpose platforms that can penetrate every industry.
- Even the core functions have yet to be fully realized, while objectives, users and form factors remain unsettled. It is therefore too early to talk about an endgame competition over value for money, price and efficiency. “I really think there still aren’t enough people in the field.”
2. Fundraising Is Essential, but Valuation Cannot Be Detached from Value
- Facing the industry’s race over fundraising and valuation, 张巍 admits he “gets distracted,” but says he does not “get particularly anxious or troubled”; LimX Dynamics remains relatively conservative on fundraising and PR, insisting that value and valuation should maintain a healthy ratio.
- Capital is hardly optional: embodied AI requires years of R&D investment, while future mass production and delivery will also consume substantial funding. “Capitalizing funding is an important capability.”
- His red line is that fundraising must serve a clear development objective—not “failing to figure out what you want to do, then raising a large sum first and figuring it out later.” Market heat creates favorable conditions, but cannot replace product and commercial judgment.
3. Industry Needs Embodied AI, but Bipedal Humanoids Need Not Enter Factories
- 张巍 draws a deliberate distinction between his industry view and LimX Dynamics’ own choice: industrial settings will certainly see extensive embodied-AI applications, but the company’s bipedal humanoid is not positioned as a factory-efficiency tool. Its slogan is “Serve people, not process.”
- The issue is not whether a humanoid can do the work, but whether it is the optimal solution for maximum efficiency and cost performance. “Factories are designed for machines,” while the human form is suited to serving people in human environments.
- Robotic arms, new manipulation systems with tactile sensing and other purpose-built machines remain better suited to replacing monotonous, complex, dangerous and hazardous work. LimX Dynamics is excluding bipedal humanoids from factories, not rejecting industrial robots.
4. Humanoids Maximize Task Count; Purpose-Built Machines Maximize Efficiency
- In 张巍’s framework, intelligence is the destination and form factor is merely the terminal that carries it: full humanoids, “wheelchair-type robots” with an upper body mounted on wheels, quadrupeds, bipeds and two-wheeled bipeds each correspond to different objective functions.
- His formula for humanoids is “maximize number of tasks over form factor”—cover the largest possible range of human tasks with a single configuration. Once 3 or 4 task categories are listed, a full humanoid often becomes the only viable solution.
- The clearest test is going downstairs to collect a package: the robot must navigate stairs and narrow doors, reach the shelf, retrieve the package and bring it back. The human environment has already written legs, arms and body dimensions into the task constraints.
- Purpose-built machines, by contrast, maximize efficiency within a defined task. The TRON product line uses a single base with different configurations to cover inspection, logistics, crossing steps and dual-arm manipulation beyond the humanoid form.
5. The First Step for World Models Is Demystification
- 张巍 first strips “world” back to a modifier of “model”: every physical or nonphysical model humans have used can be viewed as a world model at a different scale, with a different degree of openness and set of observable variables.
- Strictly speaking, a world model uses the system’s current state and the actions that may affect it to predict the state over a future period, along with the observations corresponding to that state.
- Because force-sensing data remain scarce in embodied AI, observations are currently mainly visual. A typical task is to predict the video a robot will see while performing an action, which naturally connects world models with video generation.
- Generating future video alone cannot complete a task, so the field is moving toward World Action Models: models that predict both how a task evolves in video and the corresponding action, extending the traditional VLA policy.
6. World Models’ Popularity Reflects VLA’s Scaling Bottleneck
- 张巍 does not view this as a sudden major breakthrough, but as an extension of the traditional VLA paradigm: VLMs were previously used as the backbone, while the new approach experiments with video-generation models and related technologies.
- Time-series video expresses how the physical world changes more effectively than static frames. At the same time, ego-centric videos of human activity are easy to collect, and the internet contains a huge archive of historical video, giving the field far more data than real robots can provide.
- LimX Dynamics began exploring video-data pretraining for manipulation models around mid-2024 and released VGM, or Video Generated Motion, in early 2025; it also has a CoRL paper titled GVF-Tape.
- His cooler assessment is: “It probably shouldn’t count as a major breakthrough. Otherwise, so many people wouldn’t be able to do it.” What matters is whether the approach can unlock a scaling law for embodied data, not the label itself.
7. Physical Equations Compress Data; Alignment Is the Real Problem
- Asked about injecting equations for gravity, friction, fluids, deformable bodies, temperature, electricity, mass and magnetism into models, 张巍 first preserves his uncertainty: “This question cannot currently be falsified, and it is also prone to debate.”
- His unified view of data is that neural networks and physical laws are both forms of data compression. “Newton’s laws are also a compression and representation of all motion data”; adding equations means adding historical motion information in a highly compressed form.
- The value rests on 2 questions: does the approach provide genuinely new information, and even if it does, can the model effectively align that representation with the representations already present in the world model? The first is valid in principle; the second could nullify the incremental information or even make it counterproductive.
- In practice, researchers often use physical equations to build simulations and then generate training samples that are easier to align with visual data. The entire sim-to-real process is about handling this alignment, so “useful” does not mean “easy to use well.”
8. Embodied Data Is Becoming a Raw-Materials Industry
- 张巍 directly compares model production with manufacturing: “Data is essentially a raw-materials industry, training is a production line, and the final product is a model.”
- Raw data passes through collection, inspection, preprocessing and training before becoming a model. The ability to reliably supply visual, force and other modalities of raw material therefore determines what the production line can make.
- This explains the surge of startups focused on data collection and processing. But having more data categories does not mean that pouring them into a single training pool will naturally produce general intelligence; the relationship between data types still has to be identified task by task.
9. General Capability Must Grow from a Scenario Flywheel
- 张巍 believes embodied AI cannot copy the foundation-model sequence of “build a general model first, then build applications.” Language is a general-purpose modality: a lawyer’s letter and an ordinary document can contribute to each other’s training value, but robot skills may not work that way.
- He uses the sharpest possible counterexample to explain data correlation: “Put driving, autonomous-driving and egg-peeling data into the same training run,” and today that could even create problems.
- Without understanding the relationship between cross-skill data, stacking data and waiting for emergence is “carving a boat to mark where the sword fell.” LimX Dynamics first builds a degree of general foundation, then enters vertical scenarios to collect and deploy data, with those scenarios feeding the general model.
- The approach is summarized as a “flywheel between general models and scenario data”: generality is not a product that must be completed before launch, but a capability that gradually grows through repeated validation in real scenarios.
10. The Unit of Robot Deployment Is the Skill, Not an Infinitely Large Total Model
- The host’s question was practical: if one model must cover every physical task, it could be larger than a language model with hundreds of billions of parameters, while also needing to execute continuous actions in real time at low power.
- 张巍’s answer is, “I have always felt that it is one skill at a time.” Driving is one skill and peeling an egg is another; humans have brains but may not know how to drive, and robots likewise do not need a single model that can do everything.
- The current opportunity is pretraining, post-training and deployment for individual skills. What has not yet broken even is the economics between the commercial value created by a skill and the data cost required to create it.
- Autonomous driving is itself an embodied skill, and the world model it requires may not be the same as the one needed to peel an egg. An all-in-one model covering both is conceivable, but the commercial priority is to find skills that can monetize independently.
11. A World Model Only Needs to Predict the Future, Not Reproduce Human Semantics
- 张巍’s stricter definition approximates a Markov process: an agent uses observations—which may consist only of pixels or may also include touch—to progressively recover the laws relevant to action and predict how the world will evolve after an action.
- The model does not need to recover the state of every molecule in a room. All physical laws are lower-dimensional representations of the microscopic world and all have boundaries of applicability. If predictions align with observable outcomes, the model is valid at that scale.
- The host used raw eggs, boiled eggs and iron eggs to probe semantic understanding: humans first determine “what it is” and then decide whether to peel it. 张巍’s response is that a machine can form its own embedding; “we have to allow AI to have its own semantic expression.”
- It therefore does not need to pass through a semantic layer defined by humans. “Whatever—if it can predict the future, that is enough.” What matters is the result of the action, not whether the internal concepts match ours.
12. A Model Is Not a Brain; the Brain Is an Agentic OS
- 张巍 draws a firm boundary around the “robot brain”: “I don’t think a model is a brain. A model is not a brain; a brain is not a model. A brain is an operating system.”
- This brain is an Agentic OS: it manages memory, storage, thought and task state, and calls LLM, VLM, VLA and other tools based on the objective. Stacking operational data can train skills, but does not automatically produce a brain.
- He uses OpenClaw to explain the counterintuitive relationship: OpenClaw can be viewed as a brain, but its capabilities depend on which large language model it uses underneath. The model is a capability called by the brain, not the whole brain.
- “Skill” also does not mean only atomic actions such as picking something up, putting it down or twisting a bottle cap. Even if an entire autonomous-driving system were ultimately implemented by one model, it would still be a skill—albeit a very broad one—not equivalent to the human brain.
13. Three Layers Handle Movement, Skills and Thought
- LimX Dynamics divides the system into 3 layers: the “cerebellar foundation model” at the bottom, Human Brain VLA as System 1 in the middle, and CoSA Agentic OS as System 2 at the top.
- The cerebellum turns intent into movement; Human Brain VLA binds vision, environment, task and motion into executable higher-order skills; Agentic OS handles reasoning, decision-making and skill composition.
- 张巍 uses the example of someone who is “paralyzed but extremely smart” to illustrate the separation between brain and movement: he has a brain but cannot move; enabling one skill is like giving him back the ability to reach out and pick up a cup.
- Going downstairs to buy coffee requires higher-level orchestration: deciding whether to go, standing up, opening the door, locating and navigating, calling GPS and adjusting to the environment. GPS should not be retrained into a single model simply for the sake of using GPS.
14. “Understanding” Can Be Defined by Behavioral Consistency, Not Internal Symbolic Consistency
- The host challenged the idea that an LLM, as a probabilistic prediction model, naturally proves that it truly understands semantics. 张巍 responded: “What is understanding?”
- His operational definition is that if a system can predict, and the result is similar or consistent with what a person does next, it can be called understanding; there is no need to prove that it uses the same embedding as a human.
- On whether language is necessary for thought, he limits the conclusion to the “current technological paradigm”: LLMs carry representations of thought accumulated by humans through language and the internet, but other intelligent agents may use their own embeddings and System 2.
- He also offers no definitive answer on whether machines use far more data than humans. The human brain inherits a “pretrain model” built through a long evolutionary process, while large models ingest the union of all publicly available human information; the 2 are difficult to compare directly.
15. Products Should Be Organized Around Users and Tasks, Not Software and Hardware
- The host summarized LimX Dynamics as humanoids, modular purpose-built machines, skills and an OS. 张巍 instead emphasizes that the product is “fundamentally an integration of software and hardware,” with the configuration determined by which users it serves.
- TRON is a base for innovation and scenario deployment. It can explore bipedal, two-wheeled bipedal and dual-arm platforms to cover inspection, logistics and mobility across steps, with TRON 2 placing particular emphasis on dual-arm payload and ease of use.
- Users can take the platform directly into a vertical scenario to collect data, train and validate, without first building a new machine before the commercial value is clear. Once validation succeeds, customers can continue using LimX Dynamics’ platform or manufacture their own.
- This addresses the most time-consuming phase of early deployment: POC and PMF. If proof of concept takes too long, a company may miss the critical market window before settling on a hardware configuration.
16. Open Source Should Focus on the Model Production Line, Not Parameters
- LimX Dynamics open-sourced Flux VLA Engine, but did not make the VLA model itself the centerpiece. 张巍 believes that selectively opening a limited number of parameters can “help with fundraising” today, but may not actually be usable by customers.
- The company’s approach is “give people the means to fish rather than the fish”: open the methodology and overall architecture for training a VLA foundation model, effectively handing the model production line to scenario owners.
- Vertical users own their data and should also own the models generated from that data. LimX Dynamics provides the hardware platform, interfaces, open-source architecture and necessary support, allowing scenario owners to build their own data flywheel.
17. Commercialization Will Advance from Performance and “Talking Without Touching”
- The core product logic for a humanoid is to keep one unified body and expand functionality by continuously adding apps rather than changing the hardware configuration. Research, performance, hosting, guidance, carrying water and collecting packages can all be viewed as apps added over time.
- 张巍 divides applications into no interaction, weak interaction and strong interaction. Performance falls into the no-interaction category, but it is no longer merely sell-in: customers can make money with the robots, which are already being used and delivering new experiential value.
- The next step is weak interaction that “talks but does not touch”: commercial services, guidance, sales assistance and advisory work. This is a mobile advisor equipped with language capabilities, some emotional value and a new user experience, without necessarily needing to alter the physical world first.
- Strong interaction ultimately comes down to the economics of an individual skill: can the commercial value it creates cover pretraining, post-training and data costs? Finding one vertical where the numbers break even could create a closed loop.
18. Drones Show How Performance Can Be the First Link in a Cost-Reduction Flywheel
- The host revisited drone investment from 2015—2017: after professional filming, one of the earliest applications to scale was drone performance, which had initially been underestimated.
- Demand for performances expanded fleets from dozens of units to thousands and even tens of thousands, forcing advances in flight control and coordination while lowering hardware costs through a high-volume commercial model; EVs then pushed down sensor costs further.
- Only as models, control systems and obstacle avoidance continued to improve could drones become To C products requiring no professional operator. The end state that was imaginable 12 years ago may in reality take another 10-plus years.
- 张巍 believes humanoids may follow a similar path: the ceiling for performance is limited, but performance is meaningful if it is the first of many apps. The ultimate destination remains commercial spaces and homes, which are full of diverse, low-frequency tasks that do not justify a dedicated machine for each one.
19. A True Humanoid Foundation Model Must Move Beyond Action Replay
- Many current displays of dancing and somersaults are essentially one policy playing at a time: actions are trained or recorded in advance and then replayed. They demonstrate control, but do not show a brain directing the body in real time.
- Real work requires a cerebellar Foundation Model: the upper layer provides a reference trajectory or action prompt, and the lower layer generates and executes it in real time. The system should not need another week of training just to change the way it grasps a cup.
- The host compared this with the “sensory integration training” children undergo at around 2 to 5 years old: grasping, crawling and the coordination of gross and fine motor skills first establish a physical foundation, after which higher-level intent has an execution system it can call on.
20. Market Heat Will Fade; Organization and Attitudes Toward Failure Determine Who Survives
- The host raised an investor’s contrarian concern: a concentration of flagship companies going public could signal that a boom is nearing its peak. 张巍 believes embodied AI may be more like the EV industry when XPeng went public, with public markets instead creating a capital-concentration effect that allows listed companies to catch a new wave.
- His bottom-line view is that embodied AI has “a higher ceiling than foundation models, and a higher floor”: general-purpose humanoids offer enormous upside, while even if the main route stalls, a company can still move into vertical markets rather than losing all value because it fell behind on one generation of models.
- For academic founders, 张巍 outlines 5 stages or transformations: from academia to technology, technology to engineering, engineering to product, and product to commercialization. Each transition requires a new choice of people, organization and objectives, rather than simply requiring the founder to change. He also emphasizes that the fit between talent and role has a time horizon.
- 李翔 said his most important shift in perspective was realizing that “startups can die.” Accepting the possibility of failure leads to the conclusion: “If you cannot afford to lose, you cannot win.”