Robotics Investor: Read the Papers; Commercial Scale Is Far Off
Summary
- Robotics valuations have outrun the products, but overheating does not necessarily mean a terminal bubble. 刘一鸣 notes that 1X NEO has been criticized for relying on teleoperation; he also relays a friend’s account that a Figure Demo was filmed more than 10 times before succeeding once. Both point to Demo stability, success rates and generalization falling far short of the marketing. 邱谆 believes the robot body, brain, cerebellum, simulation and world models are converging, but current capabilities still do not justify the valuations.
- Embodied intelligence may go through both a GPT-3 moment and a ChatGPT moment, rather than being ignited instantly by a single breakout product. The first phase is when data can finally train a model whose capabilities emerge as parameter scale increases; the second comes only after post-training brings it to a usable level. 邱谆 expects a GPT-3-style emergence in roughly 2-3 years if history is a guide, followed by the first market-accepted generalized use case another 2-3 years later—but “there will not be robots running all over the streets.”
- The U.S. and China are not in a simple win-or-lose contest: the U.S. is stronger in foundation models, while China leads in hardware iteration and more open data environments. U.S. efforts such as π and Skild lean more toward software and academia; Shenzhen robotics hardware can “iterate three times in a single day,” and Chinese factories are more likely to trade deployment for data. The ultimate advantage will depend on vertical integration: training the model while closing the software-hardware loop in real-world settings.
- The key to commercialization is not how many units are sold initially, but whether the business can complete the cold-start loop of deployment, data collection, training, higher success rates and broader deployment. Christine believes U.S. labor costs and willingness to pay create the highest ROI, while China may start the data flywheel first. 邱谆 is more cautious: “Whether China goes first or the U.S. goes first, I’m not particularly able to predict that yet.” Which comes first—a robotics DeepSeek or a robotics OpenAI—remains an open question.
- Embodied intelligence and intelligent manufacturing follow two different investment logics; the dividing line is whether a company must rely on large volumes of human data to generalize. AGVs, robotic arms and robot vacuums can become more intelligent while remaining specialized equipment. True embodiment will likely be close to humanoid because teleoperation, motion capture, demonstration and video data all come from humans. 邱谆 uses “three arms” as the boundary: they may be more efficient from an engineering standpoint, but are unsuitable for general-purpose embodiment because “there is no way to collect data from a person with three arms.”
- When capital rotates rapidly from one theme to another, papers provide a more forward-looking signal than fund flows. RT-2 can drive investment into VLA, but hotter funding for dexterous hands does not mean the technology has already broken through. 邱谆 even suggests using Transformer variants, diffusion models and compute requirements to assess Nvidia and the broader market, rather than inferring technology from prices. “Invest from the technology outward, not the other way around.”
- Near-term revenue will still come mainly from structured B2B use cases, and many orders that look like humanoid-robot demand are actually better suited to specialized automation. Retail replenishment and inventory checks, logistics depalletizing, factory screwdriving and external-warehouse inspection all have demand, but most current examples remain unstable demos with “very high failure rates.” Suppliers are planning annual capacity of 100,000 to 1 million units, while Goldman Sachs forecasts only 1.38 million humanoid-robot shipments globally in 2035—an obvious mismatch between orders and capacity.
- The lack of touch, hallucinations, the Sim-to-Real gap and human-robot safety together mean household robots remain a long way off. 邱谆 does not deny these constraints, but believes engineering will find shortcuts, just as aircraft achieved flight without replicating birds’ flapping wings. The real baseline is proving safety with data. If 1X has remote operators watch homes in real time, it has not only failed to clear the “Trojan horse” cold start, but has pushed the privacy issue “to another level.”
Deep dive
1. Valuations Have Priced in the Demos, but the Technology Has Not Lost Its Direction
刘一鸣 opened with two sharp contrasts: 1X markets NEO as a robot that can enter the home, yet it was quickly criticized for relying on remote operation. Meanwhile, suppliers are aggressively expanding capacity even though major real orders have yet to materialize. The market wants generalized intelligence; what it often sees is “a human behind the robot.”
邱谆 breaks overheating into two dimensions. Based on current product capabilities, Demo stability, success rates, scalability and generalization “cannot be matched to today’s valuations.” But based on the potential long-term market for physical AI, venture capital’s early positioning could still be absorbed by the size of the eventual opportunity.
邱谆’s framework is that “the eve of a major explosion is always overheated.” He acknowledges that some companies are more focused on fundraising and marketing, with videos potentially using CGI, speed-ups or repeated filming. Other teams are steadily publishing papers and introducing new architectures and models. The investment task is not to deny the heat, but to find opportunities with clear technical trajectories within it.
2. Robotics Is Still in Its BERT Phase; the GPT Moment Will Likely Come Twice
邱谆 compares the current stage with the BERT era: the Transformer path is visible, and approaches such as VLA, RT-2 and π0 have given the industry “a general sense that this is the direction to pursue.” But no model has yet reached a sufficient level in both parameter scale and performance.
The first milestone would be GPT-3-style emergence: teleoperation, motion capture, demonstrations, video, simulation and other data finally becoming usable by a sufficiently large model, much as the 175B model in 2020 trained on data that had previously been difficult to exploit. “We are very much looking forward to an emergence.”
The second milestone would be more like ChatGPT: post-training, RLHF and related methods turning unstable capabilities into something genuinely usable. 邱谆 notes that GPT-3 was still not accurate enough when it first emerged; robots may likewise show surprising capabilities first, then face a long productization cycle.
Christine offers a more concrete technical test: a robot should understand through language and vision “go to the kitchen, get a cup, pour water and put it back on the table,” decompose the task into a long action chain, and move across L0, L1 and even part of L3 autonomy rather than execute a script. Consumer adoption may look more like the iPhone: slow at first, then suddenly accelerating once the data and use cases are in place.
3. The U.S. Leads in Foundation Models, China in Hardware Velocity; the Winner Will Integrate Both
邱谆 sees a division of labor across the stack: the U.S. leads in using foundation models to drive embodiment, with π, Skild and 李飞飞’s team carrying a strong academic flavor. China is better at robot bodies, components and supply-chain iteration. The two ecosystems are not separate: Chinese companies closely track U.S. models, while U.S. teams still depend on China’s mature supply chain.
The most representative line Christine heard in Shenzhen was that robotics hardware “can even iterate three times a day.” Silicon Valley has neither the capability nor the nerve to move at that speed. But hardware velocity alone is not a complete capability; the body and the intelligence still have to be integrated in real-world settings.
The shared exception is Tesla. It has software, models and FSD capabilities, while also spending years learning from Chinese manufacturing. Christine sees Tesla as the closest current example of software-hardware integration. 邱谆 cautions that if foundation models ultimately become dominant, Google’s investment from DeepMind through Gemini could give it a stronger position than Tesla.
4. The Data Loop, Not the First Robot Launch, Determines Who Commercializes First
邱谆 defines commercialization as a loop: obtain data in a real setting, train the model, improve the success rate and only then expand. Delivering too early without enough data leads to poor performance and customer churn, eventually returning the company to the starting point of having no users and therefore no data.
China’s potential advantage is more open operating environments. Christine cites a robot piloting on a Mercedes production line that could only reproduce motions inside a tent-like “black box” to avoid sensitive production data. If a Chinese factory accepts the robot’s capabilities, it may allow deployment on the line and treat the resulting data as part of the partnership.
The U.S. advantage is ROI. Labor substitution in logistics, elder care and other labor-short sectors is highly valuable, and customers have stronger purchasing power. Christine identifies 3 conditions for scale: the business cannot remain stuck in pilots, the data flywheel must start turning, and the highest-value deployment markets may ultimately be in the U.S.
Both guests therefore leave the conclusion open. China may start the data flywheel through its operating environments and supply chain; the U.S. may clear the threshold through foundation models and high ROI. “Will a robotics DeepSeek appear first,” or will a robotics OpenAI emerge first? There is still no answer.
5. Supply-Chain Expansion Is Both a Land Grab and an Order-Unvalidated Bubble
刘一鸣 cites Goldman Sachs research showing that robotics suppliers are broadly planning annual capacity of 100,000 to 1 million units while having “received almost no actual orders.” Goldman’s forecast for global humanoid-robot shipments in 2035 is only 1.38 million units. Planned capacity has already moved well ahead of visible demand.
Christine interprets the behavior as supply-chain FOMO. Electric vehicles moved rapidly into large-scale capacity, encouraging manufacturers to “take supply-chain capacity to win orders” rather than build capacity after orders arrive. Even sensor companies have begun telling physical-AI stories; the substance is early positioning.
The risk is that robot designs have not yet stabilized. Christine says Optimus may overhaul its hardware design again around July or August, with related orders put on hold. Even suppliers that correctly identify the direction still face the trial-and-error costs of changing specifications, high rework rates and repeated mass-production adjustments.
6. The Need for Human Data Separates Embodied Intelligence from Advanced Manufacturing
邱谆 insists that most projects before ChatGPT should be called robotics, advanced manufacturing or specialized equipment—not embodied intelligence. The simplest test is whether the system needs data: if it does not require large-scale human behavior data to train the model, an intelligent robotic arm is still specialized equipment.
A true embodied route will likely be close to humanoid because the available data is centered on humans: first-person and third-person views, teleoperation, motion capture and demonstrations all follow the same template. Legs can be replaced by wheels, like “a person sitting in a wheelchair,” but a standalone robotic arm or a configuration designed for one task does not have the same data foundation.
The three-arm example is 邱谆’s sharpest counterexample. Specialized manufacturing equipment can use 3 arms for efficiency, but general-purpose embodiment will not, because “I also cannot find a person who can control 3 arms at the same time.” Configuration is not an aesthetic choice; it is defined in reverse by the training data.
7. The Upper-Body-versus-Lower-Body Debate Will Ultimately Give Way to Full-Stack Vertical Integration
邱谆 rejects dividing the opportunity simply into legs, waists and hands: “Put simply, it is the whole body.” The brain, control algorithms, actuators and body must eventually form a brain-driven vertically integrated system. Companies focused only on the lower body or control algorithms can become suppliers, but they must stay aligned with the upper-layer model or risk being displaced by upper-layer software.
刘一鸣 adds a valuation contrast: before 2023, Unitree’s valuation was at one point only half of 智元’s, or even lower. 智元 received a higher price for its more full-stack, software-oriented narrative. 邱谆 also cautions that Unitree’s current research shipments do not prove its final commercial path; customers may still switch configurations after the research phase ends.
Christine is explicit about priorities for early-stage investment: “The focus must be the brain,” especially end-to-end algorithms, data acquisition, pre-training efficiency and L2 capabilities from perception through planning. She believes current robots have not even reached the perception level of L4 autonomous driving in 2018 or 2019—if a person suddenly jumps in front of one, it may still be unable to stop.
Dexterous hands are another software-hardware intersection worth backing. Whether the design uses 3 fingers or 5 depends on the target task. Manipulation is constrained by hand mechanics, sensors and supply-chain maturity, but is ultimately controlled by the brain. Commercial value may not accrue only to the complete robot; critical components can also occupy high-value positions.
8. Investment Themes Drift; Technical Breakthroughs Are the More Reliable Leading Signal
The guests review several narrative cycles. From 2022 to 2023, Optimus went from engineering prototype to Demo machine in less than a year, while Figure, 1X and Agility drew attention to complete-robot stories. Last year, the market debated general-purpose embodiment versus scenario-specific robots. This year, the stack has been split apart and the data bottleneck has become explicit.
邱谆 relays that Physical Intelligence apparently said in a public talk during the first half of this year that data was “simply far too scarce.” Capital then began separately seeking data, model, control, body and simulation teams. 邱谆 sees this mainly as money moving between windows in the technology stack, not evidence that every rotation corresponds to a real milestone.
The causal direction cannot be reversed. A breakthrough such as RT-2 can bring investment into VLA, but a sudden inflow of capital into dexterous hands is not proof that dexterous hands have broken through. “A technical breakthrough will definitely shift investment hotspots, but the reverse is not necessarily true.”
邱谆 therefore wants investors to read the papers. During the BERT era, almost nobody paid attention to GPT-1 or GPT-2; the market was impressed only when GPT-3 trained a 175B model. The valuable judgment is to identify the trajectory from the papers during the GPT-1 and GPT-2 stages, rather than waiting for products and valuations to confirm it.
9. Robotics Investment Is Now Tied to Nvidia and the Broader AI Bubble
邱谆’s advice to public-market investors is not to track fund flows, but to study new Transformer architectures, how they combine with diffusion models and their compute requirements, then decide whether Nvidia still stands to benefit. Papers affect not only robotics names, but also whether AI capex can continue.
Having lived through the 2000 dot-com bubble, he sees the current risk more broadly. Language models and multimodal systems are also running into problems training on new data and hitting limits in scaling laws, creating demand for a new model. He tentatively calls it something like XT or XPT. If such a model fails to emerge, equity valuations that depend on continued AI progress will also come under pressure.
Without that leap, robots will not get a more reliable brain, and valuations in the equity market that depend on continued exponential AI progress will come under pressure as well. 邱谆’s conclusion remains the same: “Invest from the technology outward, not the other way around.”
10. The Lack of Touch Is a Real Constraint, but Not Proof That Embodied Intelligence Is Impossible
刘一鸣 cites Brooks’s argument that current training relies heavily on vision, while dexterous manipulation also requires touch, force feedback and skin deformation data. There is almost no historical stockpile of such data, making genuine dexterity and generalization difficult to achieve over the next 2-3 years.
邱谆 “strongly agrees” with the problem itself, but rejects turning it directly into a negative investment conclusion. Aircraft achieved flight without copying birds’ flapping wings; robots may likewise use engineering shortcuts instead of fully replicating the human sensory system. “These are precisely the problems that need to be tackled now.”
His core judgment remains vision first. Vision and language already have decades of internet data, while new tactile sensors must start collecting from zero. Models may add touch and force feedback, but will likely first improve training, denoising and precision control to fully feed existing VLA data into the model.
邱谆 distills 3 valid warnings from Brooks’s paper: data is too expensive, data structures are scarce and the final form of a robot foundation model has not yet appeared. He rejects the claim that “robots will never learn,” but acknowledges that these 3 points are the most substantive technical and investment risks today.
11. 1X’s Problem Is Not Teleoperation Itself, but the Lack of a Trojan Horse into the Home
Household safety must first be demonstrated with data. 邱谆 draws an autonomous-driving analogy: vehicles spent 3-4 years under human supervision, with intervention statistics and gradually expanded operating conditions, before earning the right to tell regulators they could drive without humans. Robots living directly alongside older people and children need a comparable safety record; that is the “baseline.”
The weight of bipedal robots magnifies the risk. 邱谆 says Unitree’s early normal-height models were too heavy and required several researchers to support them, prompting the later introduction of shorter models. Algorithms, materials and components may eventually reduce weight, but current systems have not even solved task-level success rates.
邱谆 suspects 1X may want to collect household data through teleoperation and believes the broad direction is not wrong, but says it is “a little rushed.” He notes that Waymo still operates at roughly 1 person supervising 5 cars. Monitoring is feasible in structured B2B environments; consumers are unlikely to accept a camera in the home with a human always watching from behind it.
Tesla solved the cold start by first selling an electric car with standalone value, then offering FSD as an upgrade—the car was the “Trojan horse.” 1X lacks an equivalent vehicle and would effectively be asking consumers to buy only an autonomous-driving function from a new brand on day one. 邱谆’s view is that nobody would buy a car that could only drive itself on the first day.
12. Near-Term Orders Are Hiding in Automation Gaps, but Most Are Not General-Purpose Embodiment
Christine’s favored near-term settings are industrial and retail. Replenishment, unloading, counting and inventory checks in the U.S. and Japan represent strong demand, and factories are experimenting as well. But “everything is still at the Demo stage,” and the Demos remain unstable with high failure rates.
Christine describes a logistics depalletizing motion that she believes Amazon should already have adopted: boxes enter with barcodes facing random directions, and the robot uses vision to identify the orientation and turn each box the right way. She agrees that this looks more like repetitive work or specialized equipment than embodied intelligence.
Screwdriving also requires a distinction. A fixed position can be handled by a robotic arm, but positions, alignment and torque may vary on an automotive production line, requiring stronger generalization. 邱谆 believes BYD may have this kind of demand, yet the industry is still collecting data through motion capture, teleoperation and demonstrations, with no company having clearly solved the problem.
The real opportunity will enter through automation gaps. At Foxconn, 邱谆 saw the internal warehouse already 100% automated, while the external warehouse still required 2-3 people to pull boxes, inspect them, close cartons and apply bands. He believes embodiment’s target is precisely this work that conveyors, AGVs and collaborative robots cannot fully cover.
13. Hallucinations Are Not a Robotics-Specific Problem but a Weakness in the Entire AI Architecture
刘一鸣’s challenge is that Transformers cannot eliminate hallucinations even in language, while industrial production lines demand reliability approaching “many nines.” If a black-box model generates an incorrect action, it can cause not only line stoppages but also physical collisions with people.
邱谆 fully accepts the risk. AI outperforms traditional algorithms because it learns from data rather than relying on line-by-line rules, but the same mechanism makes the process difficult to explain. Embodied systems are still generating the next token; the output has simply shifted from text to positions and poses. “If your generated language has hallucinations, how could your generated poses possibly have none?”
Agent and Manus-style multi-step calls currently use rules, reasoning and repeated calls to correct the success rate of a single large-model inference, but they add compute and execution time. The real advance still requires a leap in AI itself, followed by joint gains from pre-training, post-training, fine-tuning, self-supervision and the Agent layer to improve end-to-end execution rates.
14. World Models and Sim-to-Real Can Supplement Data, but Cannot Replace Real Robots
Christine believes a world model designed for robots “may be promising” because, as a generative model, it could improve accuracy through engineering techniques. But 李飞飞’s direction will not solve the entire problem on its own; it will ultimately need to be combined with VLA, real-world data and other model modules.
She is particularly interested in the intersection of world models and games. Users interacting in a 3D world effectively generate continuously labeled data, which can then be packaged for robot training. Games can also generate cash flow of their own, “just as the gaming industry helped nurture Nvidia’s GPUs back then.”
邱谆 is more restrained on Sim-to-Real. The gap between simulation and reality is persistent, and industrial digital twins have faced the same issue. One or several high-quality real-robot data points may be worth more than large volumes of simulation data. Simulation still requires humans in the loop to calibrate and filter noise, so its total cost may approach or even exceed real-world collection.
Simulation, synthetic data, Video Gen and Self-Play are therefore supplements to the data pipeline, not shortcuts that solve the problem through unlimited generation. “Once everything truly converges, the most useful thing may still be a piece of real data.” The remaining questions are how to define the specification and obtain the data cheaply.
15. The First Generalized Use Case May Arrive in 5 Years; Household Adoption Is Farther Away
The hardware infrastructure already exists. Reducers, motors, sensors and mechanical structures are progressing more linearly, while software may allow ordinary “off-the-shelf” components to deliver new value. 邱谆 cites Velodyne as an example: the hardware did not suddenly change; software finally made it possible to incorporate its data into training.
Robots still need more reliable bodies, and their specifications have not yet settled. Christine expects the next cycle to focus on robustness, resilience and rework rates, with better hardware iterations possible next year. 邱谆 believes next year may bring clearer convergence around the use cases suited to world models.
邱谆’s 5-year timeline is roughly 2-3 years to GPT-3-style data emergence, followed by another 2-3 years to reach ChatGPT-style usability. By then, there may be only a first generalized use case rather than broad deployment. 刘一鸣 suggests household robots may still be 13 years away; the guests do not confirm a precise year, but agree that household adoption will be slower.
Christine expects adoption to proceed from structured B2B production settings, to continuous work in restaurants and similar environments under human control, and then to the home once trust has been established. 邱谆 adds that startups will not always produce dazzling Demos; they must also handle the “dirty work” of data collection, reliability and scaling. The most eye-catching results may first appear in research papers, but research breakthroughs may not scale commercially.