Why Lingxin Tops Morgan Stanley's Humanoid Report: Founder Wang Qibin
Why Lingxin Tops Morgan Stanley's Humanoid Report: Founder Wang Qibin
Summary
- Lingxin Intelligence has just topped Morgan Stanley’s global humanoid-robotics report and entered the first tier of embodied-brain companies; founder Wang Qibin’s positioning is that, from Day One, it would be a brain company focused exclusively on “general-purpose dexterous manipulation,” solving “the crown-jewel problem in embodiment: the manipulation problem of the human hand.” In mid-April, the company released its Policy model R2 and Action-Conditional World Model, ranked first on the U.S. Mobile ALOHA leaderboard, and saw its 1,000-hour multimodal dataset on Hugging Face become the platform’s most-downloaded dataset. For the previous 8-9 months, most of its effort had gone into data-collection hardware and pipelines.
- Wang Qibin’s direct test for whether a company is genuinely training a brain or is a “sham company” is: “How much compute does it spend each year?” Lingxin’s compute bill this year is conservatively estimated at a few tens of millions of yuan and may not cover the full requirement. Real-robot data is expected to rise from 100,000 hours to 400,000-500,000 hours by midyear and 1M hours by year-end—but Wang stresses that “1M hours is only a starting point.” The true requirement could be 100M hours, and no one will monopolize the data over the next 3 years that can be reasonably foreseen.
- Lingxin’s core data advantage is a human-centric, full-modality data glove capturing joint angles and tactile signals, along with head-mounted data, rather than relying only on pure ego video or gripper data. Wang believes the precision of 3D hand pose matters far more than raw data volume or 2D images. Six months ago, the same grasping data cost 2-3 yuan in China versus $2-3 in the U.S.; domestic real-data collection costs roughly several dozen yuan to RMB100 per hour. With lower data costs and faster scaling, China “may be able to catch up on models,” much as DeepSeek R1 caught up under compute constraints.
- His midpoint estimate for the industry cycle is roughly 7 years: the first elimination round has not yet begun, but the fundraising arms race and Matthew effect are already under way. He believes the embodied market could grow to 10x, tens of times, or even 100x the size of autonomous driving. Over 5-10 years, some new companies could far exceed the previous industry peak, while companies that represented the last peak could become the low point of this cycle. He also makes clear that valuation cannot fully assess a company.
- The previous generation of automation-robotics companies will probably struggle to win this wave, a judgment Wang offers as both a BlackBerry veteran and someone who experienced the innovator’s dilemma firsthand. The shift from rule-based to learning-based systems is a wholesale change in technology and talent. The eventual leaders in warehouse logistics and autonomous delivery were all startups. The BlackBerry lesson is that a new species “does not emerge from the old world”; it comes from a new world and delivers a dimensionality-reduction attack on the old one.
- “Pure model companies” are a “false proposition for the next 5 years” in embodiment: for now, hardware and software must be tightly coupled. Lingxin writes its own software for hand-position control, velocity control, and current-loop force control, while designing the hardware and having partners customize it. The commercial path is To B first, in logistics and services, with To C households as the end state. The target is several hundred million yuan in sales by the end of 2026, from two lines: data-collection systems and data, plus industry solutions. Logistics deployment requires at least a 99.9% success rate, cycle times no worse than the floor for Chinese workers, and human takeover.
- Data quality is a major risk: China has roughly 50-plus data-collection facilities, but probably fewer than 10 operate efficiently enough to produce data that can actually be traded. Wang also relays that the reported 270,000 hours of data “may” contain smoke and mirrors. On companies that have raised substantial capital, claim to be building brains, but are not actually training models, his assessment is “possible,” though he is not certain. Waiting to copy the U.S. leaders after they open-source their work is “too optimistic,” and PI has not released complete model weights either.
- At the architecture level, Lingxin says it was among the earliest companies to connect two models into a closed loop: its R-series policy model takes images, language, and machine state to generate actions; the World Model evaluates the resulting state and optimizes the action, with roughly 30% corrective data added to create a feedback loop. PI’s π0.7 showed that real-world data can generalize on grippers and decompose compositional language tasks, but Wang believes the GPT-3.5 moment is still far away; the first meaningful test of model generalization will come only after the data volume steps up by year-end.
Deep dive
1. Topping Morgan Stanley’s Report: A Concentrated Release After 8-9 Months of Silence
- Wang Qibin opened by tallying the results: in mid-April, Lingxin released its Policy model R2 and Action-Conditional World Model, ranked first on the U.S. Mobile ALOHA leaderboard, and saw its 1,000-hour dataset on Hugging Face become the platform’s most-downloaded dataset. Overseas, more powerful brain companies are also releasing general-purpose dexterous-manipulation models, including Genesis—evidence that Lingxin chose the right direction from the outset.
- Why hasn’t its valuation risen more sharply, and why isn’t the “brain” label more prominent? Wei Shijie put the question directly. Wang’s answer: after the World Robot Conference last year, the company spent roughly 6 months doing little PR and instead laying out its data-collection hardware and pipeline. It went through an 8-9-month buildout. “The power of language may be the weakest force”; for a technology company, the most persuasive evidence is the model’s performance and the dataset itself.
2. Who Is Lingxin? A Day-One Brain Company Focused Only on General-Purpose Dexterous Manipulation
- The positioning is explicit: Lingxin was incorporated in September 2024, and Wang defines its mission as solving “the crown-jewel problem in embodiment: the manipulation problem of the human hand.” It was China’s first company to commit from Day One to general-purpose dexterous manipulation without making grippers or mobile bases—a positioning that was “far ahead of the market’s understanding” at the time. The company uses algorithms to drive data collection, then uses the data to deliver integrated solutions.
- The name comes from the logo, the Greek letter ψ (Psi), the 23rd letter of the alphabet, which carries meanings in psychology and physics and evokes proprioceptive intelligence. It also maps to the reinforcement-learning process of “growing up like a child.” A co-founder proposed “Zero Yuan,” which Wang rejected: “Does that mean zero-yuan shopping?” The team eventually combined “Ling” and “Xin”: dexterous manipulation is at an initial stage and will keep developing.
3. Three Capabilities of General-Purpose Dexterous Manipulation—and an Evolutionary Argument
- Wang breaks human manipulation into 3 traits: long-horizon task decomposition—the subconscious knowledge to pick up the microphone before its cable when clearing a table; precise hand-eye coordination; and rapid autonomous correction after an error. He says humans took “nearly a year of evolution” to develop these capabilities.
- His evolutionary argument is that language was the last human modality to emerge, and machines have largely learned it from text and other information; vision emerged during the Cambrian period; and action appeared before humans became Homo sapiens and underwent the longest evolutionary process. General-purpose dexterous manipulation therefore requires action, vision, and language in one system, making it an unusually complex system.
- On the human-like division in which planning belongs to the brain and hand-eye coordination to the cerebellum, Wang acknowledges that the analogy cannot be mapped cleanly. Robots still lack a perfect “brain-cerebellum” structure for fitting or generating the closest approximation to human capabilities.
4. The Gripper Ceiling: Unscrewing a Bottle of Water
- To explain the difference between a gripper and general-purpose dexterous manipulation, Wang uses the act of unscrewing a bottle cap. Humans naturally use 3 fingers to create a force-closure torque and twist it open—a human-like manipulation pattern. Lingxin wants robots not only to complete many tasks, but also to perform fine, human-like manipulation of objects such as electronic components. Grippers struggle with this class of operation.
- The later, more data-oriented point is that a gripper-only approach faces a bottleneck in sustaining diverse, sufficient data sources. The resulting model performance would also be constrained. That is one of the key differences between the gripper route and the human-data route.
5. Human Manipulation Data: An Untapped Gold Mine
- Wang’s foundational view is that human manipulation happens every day in factories, service businesses, and homes and “must be a gold mine.” But for years there was no tool to record it. When someone opens a bottle, no one has captured how many fingers were used, how much torque was applied, or what the tactile sensation was.
- This knowledge was previously transmitted largely through capabilities humans had already internalized at the genetic level. The basic problem Lingxin wants to solve is how to extract the relevant knowledge from humans—and even animals—and turn it into data that embodied intelligence can use.
6. Industry Heat Exceeds Expectations; Midpoint Cycle Estimate: 7 Years
- Responding to the claim that “China’s embodied market is highly legendary,” Wang first acknowledged that the heat has exceeded expectations: large financings and an industry “burning hot,” like “surfing, with a new wave coming in.” But he sees basic logic behind the enthusiasm. Autonomous driving is the earliest mature application of embodied intelligence, handling horizontal movement on flat surfaces and within an existing transportation system; the market is already large. Embodiment could eventually become 10x, tens of times, or even 100x larger than autonomous driving.
- There is still little consensus on the cycle length for individual applications. Wang takes a midpoint estimate of roughly 7 years. Automotive intelligence took roughly 6 years from 2018-2019 to last year, while embodiment operates in 3D space and represents a further dimensional leap. At the same time, the current cohort has high talent density, substantial funding, and strong participants. Over 5-10 years, companies could emerge that far exceed the previous industry peak, while those peak companies could become the low point of this cycle.
7. Valuation Is Not the Benchmark, but the Arms Race Is Under Way
- Asked whether Lingxin is near the first tier by valuation and where it stands on actual capability, Wang separated the questions. In China, on general-purpose dexterous manipulation, Lingxin is “without question number one,” because its accumulation did not happen overnight. But valuation cannot fully assess a company.
- He gives 2 reasons: the embodied ecosystem and supply chain are broader and more fragmented; and this cohort of companies has been around only a little over 2 years since 2023, in a race lasting roughly 7 years with several elimination rounds still ahead.
- The fundraising arms race has begun, and global financing is developing a Matthew effect: the more leading the company, the more money it receives, and the faster it receives it.
8. A 1970s-Born Career: BlackBerry—Sonos—Yunji—JD.com
- Wang describes himself as a member of the post-1970s generation. He went abroad for school in 2001, earned a Ph.D. from George Washington University in 2008, and returned to China. From 2008-2012 he was a product manager at BlackBerry, then moved into smart speakers and worked through Sonos’s product phase.
- 2018 was the dividing line: he joined Yunji as vice president of product, then moved to JD.com’s X division, where he spent the 3.5 years before starting his company working on delivery robots and autonomous vehicles.
- His historical view of “intelligence” is that technological evolution continually rewrites its meaning. In the smartphone era, a multi-sensor platform gave rise to the app ecosystem and mobile internet. In 2018, he saw robots’ mobility as a larger breakthrough in intelligence. Starting in 2023, embodied intelligence went further: model algorithms overturned the old concept of intelligence by adding human-like thinking and reasoning, plus the ability to interact with a hardware platform.
9. BlackBerry’s Failure: A Dimensionality-Reduction Attack from a New-World Species
- The most important lesson Wang learned at BlackBerry was that it was the only company in the world during his tenure to achieve a net margin above 25%; Apple was the second. That still did not prevent BlackBerry’s eventual failure. He also cites the collapse of Nokia, Motorola, and a string of other companies.
- If he could do it again, he believes BlackBerry’s fate would be difficult to change. But it had 2 advantages at the time: BBM and email. China Mobile once hoped BlackBerry would license BBM to third-party platforms, which reminded him of how iTunes moved from Mac to Windows by opening up.
- His conclusion is that BlackBerry lacked sufficient innovation at the hardware layer, but could have made better choices at the application layer. A new species does not grow within the old world; it comes from an entirely new world and delivers a dimensionality-reduction attack on the old one.
10. Why Apple Won: Not Just an Open Ecosystem, but Years of Deep Interlocking
- Wang rejects the idea that the decisive factor can be reduced to an open ecosystem. After studying the history of companies globally, he concluded that the iPhone’s success came from years of interaction among hardware, software, and business model. Apple identified a Japanese storage technology relatively early, and iTunes already existed before the App Store.
- iTunes was disruptive not merely because it made music easier to store and download, but because it changed the business model. Previously, buying an album meant buying one CD containing both the good songs and the bad ones. iTunes broke the bundle apart and charged by the individual song.
- The iPhone was therefore the product of many factors interlocking over years; no company could have claimed in advance that it was certain to build it. Winners are usually “complex systems with deep coupling and powerful synergy,” not single modules.
11. Returning to China, the Rise of Chinese Hardware, and Embodied Products That Are Still Ugly
- Returning to China was a choice with a sense of fate. Wang disliked the East Coast atmosphere of Washington, D.C., dominated by the federal government; his family and wife had already returned; and he kept Qian Mu’s Outline of National History by his bed. He missed Wuhan’s “zao can”—the city’s early breakfast—and believes history has no “if.”
- After returning, he watched Chinese hardware move from “pretty ugly” to rapid ascent. After 2010, the Xiaomi ecosystem and IoT drove consumer-electronics industrial design to catch up quickly and begin forming its own aesthetic. The foundation was a complete industrial system covering whole-device design, core components, mature processes, and rapid iteration. He admits he regrets not entering the Chinese ecosystem earlier and more deeply.
- His aesthetic judgment on embodiment is equally direct: from a consumer-electronics perspective, most embodied products remain primitive, with little aesthetic quality and a primary focus on solving functional problems. As the ecosystem iterates and strong talent enters, China has a chance to deliver both better looks and better technical aesthetics.
- He believes aesthetics require both a material foundation and a tradition. Jony Ive’s design tradition is British, while the leader of Xiaomi’s design team has a German design background. B&O pursues dazzling visual impact; Sonos emphasizes black-and-white color schemes and integration into the home.
12. The U.S.-China Debate: “America Leads in Brains, China in Hardware” Sees Only the Starting Point
- Wang pushes back on the simple frame that “the U.S. leads in brains and China leads in hardware,” arguing that it sees only the starting point of the trajectory. The U.S. does have some lead in model iteration, but China has richer access to data and more application scenarios.
- China has also gone through successive product cycles in the internet, mobile internet, and automotive industries, building strong product-management capabilities. TikTok, Douyin, and some IoT products reflect China’s strengths in reading subtle user behavior and defining products.
- Embodiment is not the simple addition of hardware, data, models, and applications. It is a deeply coupled system spanning hardware, big data, models, and application software, so it is impossible to declare a winner by looking at any single layer.
13. To B First, To C Eventually: What Kind of To C Product Can Break Through To B?
- The commercial path is clear: the medium-term vision is To C, with the home as the most general-purpose setting. But homes require skill generalization in a highly unstructured environment and are currently the hardest scenario. Lingxin is therefore entering through To B settings such as logistics and services, where generalization requirements are moderate and cycle times are lower.
- The lesson Wang takes from the history of BlackBerry and Apple is not a simple To B-versus-To C debate, but what kind of To C product can break through To B. Apple entered the enterprise through devices, showing that platform-type products can do this.
- For robots, if general-purpose capabilities can be integrated conveniently into To B systems, they may be able to penetrate from To B to To C—or from To C back into To B. Lingxin’s models are already used for grasping, scanning codes, placing objects, and opening plastic bags; fundamentally, these remain general-purpose capabilities.
14. “Will Every Company Become a Robotics Company?” Half Yes, Half No
- Wei Shijie relayed a contrarian view from another embodied-hardware founder: there will not be 2 identical B-end robots in the future; each enterprise will define a differentiated robot based on ROI, and the best technology may be opened to enterprises rather than sold as a product. Wang admits he has “not fully thought it through,” seeing both merit and limitations in the argument.
- His breakdown starts with hardware’s biggest rule: economies of scale. The market may ultimately form a multi-player oligopoly. Large B-end customers that can achieve scale may define their own hardware, but mid-market and long-tail companies will struggle to own a hardware form factor independently.
- The brain and data layer is more uncertain and more consistent with a winner-take-more dynamic: the more data enters, the faster the flywheel turns. The brain may develop broad general-purpose capabilities, while hardware may be defined by leading companies and used by mid-market and long-tail companies.
15. “Pure Models” Are a False Proposition for 5 Years: Writing All the Way Down to the Current Loop
- On the path taken by leading U.S. companies that focus only on models, Wang is unequivocal: he personally believes the proposition is false for the next 5 years because embodiment remains in a phase of tight hardware-software coupling. Without controlling the arm and fingers, a system cannot reach the global optimum.
- High-dynamic actions require deep coupling between software and hardware from visual input through system output. Only after the technology has iterated to a certain stage can the ecosystem reach an eventual decoupling.
- Lingxin designs its own hardware, has partners customize it, and writes all of its control software itself—from hand-position control and velocity control through current-loop force control. The goal is to optimize both the system and the cost.
- He expects hardware over the next 5 years to move from generalist toward general specialist. Just as compute platforms move from GPUs toward more customized, efficient, and lower-cost chips, further hardware specialization and scaling may arrive only later.
16. Why Does It Need a Wheeled Body? The Embodied Data Flywheel Is Completely Different from the Automotive One
- If 2 arms can handle manipulation, why build an entire robot? Wang answers by scenario. Desktop manipulation needs only 2 arms and 2 hands, or 1 arm. Manipulation across a vertical plane or in a circular space requires waist movement. Restocking or picking across 4-5 square meters in a supermarket requires a fast, stable base for short-range movement.
- Lingxin therefore envisions 2 configurations: a fixed upper body and a mobile form with a body component. Both are currently trained with the same model.
- The deeper issue is how to start the data flywheel. Automobiles have an installed-base market: even without autonomous-driving systems, people drive the cars, and after users buy them they still generate data for the automaker. By contrast, leading embodied companies shipped roughly 5,000 units last year; even a few tens of thousands this year would not be enough to support a data flywheel.
- Lingxin will therefore start the flywheel with human data, then deploy its own hardware in real-world scenarios so that new data generated during inference can flow back into the system.
17. Why the Previous Generation of Automation Companies Will Struggle: An Insider’s Innovator’s Dilemma
- Asked whether the previous generation of industrial-automation companies could still win by transforming, Wang’s answer is that they will probably struggle. This wave faces more complex manipulation problems; the technology is shifting from rule-based to learning-based; and the talent required by the learning paradigm is entirely different.
- Earlier robotics companies typically built deep expertise in a single setting—hotels, restaurants, warehouses, and so on. Their installed scenarios may become a burden. This wave needs a general learning paradigm to turn on the data flywheel, break through scenarios one by one, and eventually cover 10, 20, or more scenarios.
- Wang acknowledges that Wei’s counter-scenario is “possible”: if a general-purpose brain matures, incumbents could retrofit customer systems with new technology, and data feedback might arrive faster. But once a general-purpose brain solves a scenario, who controls the most valuable asset is not a simple static question.
- His firsthand evidence is that he once proposed that his former company create a completely independent innovation and investment division, but the idea was rejected. Later, at JD.com, he believed “whoever owns the scenarios can win,” only to see the innovator’s dilemma confirmed. The warehouse-logistics leaders include Hai Robotics and Geek+, while the autonomous-delivery leaders include Neolix and 9th. Alibaba, Meituan, and JD.com all fell behind the startups.
- Mature business models continually reinforce the processes and decision-making systems of large companies, making it difficult for them to step away and pursue businesses with longer innovation cycles.
18. The Founding Timeline: 6 Months Searching for Scientists, 3 Months Waiting for the Name
- Wang rejects the view that companies with conviction appeared in 2023 and those founded in 2024 were merely following the trend. He had largely decided to start a company by the end of 2023, then spent roughly 6 months assembling the team and finding scientists.
- Lingxin received investment from Hillhouse and BlueRun in May 2024 and began operating in June. But an institution in Beijing had already registered the name “Lingxin Intelligence.” Under local rules, the name would be released only after 3 months without incorporation, so the company was not formally registered until September.
- Alignment with the co-founder also took time. Starting around September or October 2023, Wang spoke with Yang Yaodong for roughly 6 months, read his papers, and assessed their respective capabilities and fit. Only then did they settle on whether to focus long term on hands, whether to touch grippers, and how to build the initial team. They argued along the way but wanted to establish a shared belief that “slow is fast.”
19. A 1970s-80s-90s-00s Team: Assembled from Fewer Than 10 Scientists Who Could Do Manipulation
- The team brings together Wang Qibin, a product veteran from the 1970s generation; algorithm engineer Chai Xiaojie from the 1980s generation; scientist Yang Yaodong from the 1990s generation; and scientist Chen Yuanpei from the 2000s generation. The mix was deliberate, reverse-engineered from the full chain running from technology to product to market: Wang handles product and market, Yang handles foundational science, and Chai handles algorithm engineering.
- In China in 2023, very few scientists could work on dexterous manipulation; Wang estimates the number at no more than 10. He interviewed 7 or 8 people from institutions including Shanghai Jiao Tong University, Tsinghua, Peking University, and Wuhan University and felt the competition for talent firsthand. At least 6 or 7 U.S. researchers also spoke with him, including Chen Chi of Sunday Robotics, Wang Chen of Stanford, and Tan Jie of DeepMind. Some wanted to start companies in the U.S., closer to home.
- Yang returned from UCL to Peking University and was an early researcher in bimanual-manipulation simulation environments and data, applications of dexterous-manipulation data, and large-model post-training and alignment. From the end of 2022 through 2023, he also took on a nationally designated commercialization R&D project focused on general-purpose dexterous manipulation.
- 2 statements from Yang won Wang over: the place capable of producing the greatest innovation globally had already shifted from universities to outstanding startups; and papers had delivered a sense of achievement, but he wanted to create something new and meaningful in industry.
- Chen had conducted research in the labs of Fei-Fei Li and Kerri Liu. Before leaving Stanford, he published a paper on using human data to train sim-to-real manipulation capabilities in simulated environments. That work is also reflected in Lingxin’s current large-model-plus-reinforcement-learning approach.
20. Why Put the Company in China: “9 Months Later” in a Sonos Conference Room
- Wang once brought Sonos’s global CEO to China to meet Lu Qi and discuss voice collaboration. Lu had many Sonos speakers at home and said that if Sonos wanted to do anything in China, he could have his team connect immediately. But after leaving the building, Wang asked his boss when they would start. The answer was “9 months later”: first finish Amazon, then Google, with Baidu third in the pipeline.
- He came away feeling that advancing a China-specific initiative inside a global company was heavily constrained and began thinking about doing something more interesting within China’s ecosystem. He never developed the idea that he had to start a company overseas.
- His observation of U.S. and Chinese research styles is that some U.S. researchers focus more on grand questions and less on near-term deployment. Chinese researchers not only build models but also consider commercialization, so the Chinese ecosystem “reinforces” them toward greater pragmatism.
21. How a Nontechnical CEO Tests the Signal: Peer Review, Demos, and “Success Without Failure Is Not Credible”
- Lingxin uses a degree of internal peer review to determine whether a scientist’s work is solid, including mutual assessment among Yang Yaodong, Chen Yuanpei, and Professor Wen of Shanghai Jiao Tong University, followed by a look at the actual model performance.
- When choosing scientists, the first question is whether they have only published papers or have built a demo. A demo means they have taken one step further. The company then conducts background checks, examines their expertise and genuine commitment to entrepreneurship, and discusses whether core economic interests can be aligned effectively and sustainably.
- Wang interviews nearly every candidate, spending 30 minutes to 1 hour with each one. Even the most junior candidates receive 30 minutes. He repeatedly asks how they collaborate and how they would solve the hardest problems. He focuses more on their approach to problem-solving and whether they persistently use every available effort to solve the problem.
- He dislikes candidates with polished resumes and uninterrupted success, because hard problems inevitably bring technical or engineering setbacks. “Success without failure is not credible.”
22. R&D Governance: Bottom-Up During Exploration, Top-Down After Route Selection
- On the difference between scientists who can push deadlines indefinitely and production teams that must deliver on schedule, Wang says algorithmic-model research cannot be managed entirely by production logic. The earliest exploration phase is not suited to rigid time management. Only with the latest model, after the training and infrastructure pipelines were established, did the timeline become more reliable.
- He even thinks “management” may not be the most accurate word; perhaps it should be called “governance.” Lingxin is not a pure research team but a research-and-development organization, with algorithm staff working alongside data-platform, training-platform, and engineering-architecture teams in one large group.
- The organizational method is more bottom-up during exploration. Once a route is confirmed to be important, resources are allocated top-down. The company is currently driven mainly by projects, with a major data project potentially involving 70-80 people. Co-founders debate whether a research project should enter development, how much resource to commit, and whether it needs a budget, but these are tactical questions and are usually resolved quickly.
23. Compute as a Litmus Test: A Conservative Estimate of Tens of Millions of Yuan, and It May Not Be Enough
- The most quotable industry test from the episode is: “The most direct standard for telling whether a company is genuinely training a brain or is a sham company is how much compute it spends each year.”
- Lingxin’s compute spending this year is conservatively estimated at a few tens of millions of yuan and may not cover the full need, because data scaling will accelerate quickly. Wang says the compute corresponding to 1M hours of data by year-end is still being assessed.
- His evaluation is restrained: the figure is certainly not low among embodied companies, but is not particularly high by his personal standard. Alibaba, ByteDance, Xiaomi, and other major companies are training embodied brains, with more financial and compute capacity than startups and possible advantages in language models and multimodal video.
24. The Weapon Against Big Tech: 3D Hand Pose Matters Far More Than Data Volume and 2D Video
- Facing Alibaba, ByteDance, Xiaomi, and Google with its powerful language models, Wang emphasizes Lingxin’s foundational data insight: the precision of 3D hand pose matters far more than raw data volume or 2D images.
- Lingxin’s collection hardware captures human joint-angle trajectories and tactile signals, unlike approaches that train only on pure video or first-person ego video. Joint-angle trajectories make the data more precise and richer and support repeated iteration. The volume will rise from 100,000 hours to 400,000-500,000 hours by midyear and 1M hours by year-end.
- The data glove is a full-modality device, with sensors analogous to human joint angles and tactile sensors; users also wear a headband during collection. People can work while wearing the glove and collecting data, and later collect data while working in real environments.
- The pipeline is not simply handing gloves to partners. It includes data review, labeling, and processing before the data enters the training framework. The 100,000 hours of multimodal data were collected from late October last year through March this year at 3 data-collection facilities using several hundred glove sets. Wang calls it the industry’s largest hand-data dataset to date.
- Lingxin plans to expand the glove fleet to at least 1,000 sets this year. With sufficient resources, the company believes it can scale the validated performance to that level.
25. The Right to Identify Garbage: 50-Plus Collection Facilities, Fewer Than 10 Running Well
- High-quality pretraining data must capture 3 properties of human manipulation: semantic long-horizon task decomposition, multimodal hand-eye coordination, and error correction. It also needs task diversity. Beyond vision, tactile and joint-angle data are important as well.
- During collection, the industry must repeatedly test whether depth data is necessary and what cameras to use. Depth data provides vertical information such as object thickness and distance and can be collected with depth cameras, structured light, and other methods.
- The industry’s disorder has a measurable footprint: more than 50 data-collection facilities have been built nationwide, but probably fewer than 10 operate efficiently enough to produce data for trade.
- Data needs to be pulled by demand, with the buyer defining formats, modalities, and collection standards. Wang believes only foundation-model companies and leading embodied companies have enough capability to distinguish garbage from useful data. If the data itself is bad, it is garbage in, garbage out, and many participants may indeed be producing data garbage.
26. Data Costs, Data Exchanges, and a Recipe That Will Never Be Disclosed
- Six months ago, collecting the same grasping data could cost 2-3 yuan in China versus 2-3 dollars in the U.S. Current domestic real-data collection costs vary with task difficulty, from several dozen yuan to around RMB100 per hour. Costs should continue to fall across 3 areas: equipment, operations and storage, and compute.
- Wang believes that if data costs become low enough and scale ramps quickly enough, China may be able to catch up with the U.S. on models. One reference point is DeepSeek R1, which iterated rapidly despite constrained compute resources.
- On the objection that high-value data will never be sold, he says the data requirement for embodied models is far larger than expected: 1M hours is only a starting point, and the true requirement could be 100 times that, or 100M hours. Each company’s data is more like a lake in an ocean, and data exchanges will inevitably emerge at some stage. He does not believe anyone can truly monopolize the data over the next 3 years that can be reasonably foreseen.
- Data value also varies by scenario. Logistics or supermarket data is strong in its own domain, while high-precision, high-cycle-time industrial data is strong in another. Lingxin was still supplying data to leading embodied-model companies early this year; Wang says he was observing the ecosystem’s iteration as a data supplier.
- On the recipe, he says that after GPT-3 or GPT-3.5, companies stopped publishing complete technical details and data recipes. Lingxin will not disclose its actual recipe or specific hour counts to investors, only the broad data types and scale. Investors ultimately judge “how the dish tastes.”
- Wei Shijie argues that data may not be an absolute moat: a company that truly understands model training could take external data and still establish a lead. Wang agrees, saying the leverage of a brain company is high and it may not need the largest capital base.
27. Lessons from PI and an Industry-First Two-Model Closed Loop
- Lingxin has studied Physical Intelligence (PI) continuously but does not benchmark itself against PI in every respect. PI uses gripper data, while Lingxin collects high-dimensional hand data. The approach also partially resembles the Dream series—Dream Zero and Dream Dog—which focuses on the hand.
- π0.7 showed through real-machine manipulation data that a model running on real data can generalize on grippers, understand language, and decompose compositional language into robot tasks. This validates part of the real-data route, though when asked whether it had therefore fully committed to the route, Wang answered, “Partly.”
- He cautions that leading companies open-sourcing models and releasing gripper data does not mean they only work on grippers. Public releases may represent an earlier or second-earlier generation rather than the frontier. Wei says 2 or 3 companies in the industry are already working on hands.
- The first model in Lingxin’s two-model architecture is a human-like action-generation model, the R-series Action Model. It takes images, language, and historical state as inputs and outputs robot actions. The R stands for reinforcement learning, not reasoning.
- The second is the World Model, which evaluates the system state after an action is executed. If the state does not match expectations, it optimizes within the model and adds roughly 30% corrective data. The 2 models form a closed loop: the policy model generates actions, the World Model evaluates and optimizes them, and the data flows back to generate a new dataset.
- Wang says Lingxin was among the earliest companies to connect 2 models into a closed loop and propose this architecture.
28. Recruiting Through Transparency, Rejecting the “Crown Jewel”: We Are Surfing, Not Climbing
- In competing with major companies for people who have trained large models or worked on large-scale post-training, Lingxin can attract candidates who genuinely want to start companies, but “sometimes we still can’t win.” Compensation is only part of the draw. More important is whether the candidate wants to tackle an exceptionally difficult problem in embodiment and the reputation the company is gradually building.
- Recruiters tell HR that the industry sees Lingxin as a simple, transparent company. The co-founders favor low communication costs and a clear technical strategy.
- Wang rejects calling dexterous manipulation “the crown jewel,” saying the term “jewel” feels dated and unlike the language of the AI era. Nor does he think Lingxin is climbing a mountain; it is surfing. Technology comes in wave after wave, the shape of each wave keeps changing, and the company needs the ability to truly serve the wave.
- He does believe general-purpose dexterous manipulation will be a core value anchor in the embodied transformation. From human evolution, manipulation is the hardest part and the best test of algorithmic and model difficulty.
- Chen Yuanpei once suggested positioning Lingxin as an AGI company that might not even build robots in the future, but Wang did not adopt the positioning. He cites a Kantian formulation: “Look up at the brilliant starry sky above, but keep your feet on the ground.” First understand where you stand and what the path is; only then discuss a more distant vision. Once the thinking is clearer, saying it aloud becomes more persuasive.
29. Commercialization: Several Hundred Million Yuan by End-2026, with Two Hard Conditions for Scenario Selection
- The revenue structure maps to the 3 waves Wang sees: hardware has already emerged; data-collection equipment and data monetization are next; and scenario solutions follow. Lingxin aims to generate several hundred million yuan in sales by the end of 2026 through 2 lines: data-collection systems and data, plus industry solutions.
- Data customers include collection facilities and domestic technology-model companies. In logistics, Lingxin has begun POCs with leading domestic logistics companies. Industrial customers include auto-parts suppliers and 3C companies.
- On liability for a production accident, Wang says the integrator and supplier will define responsibility through agreements, and someone will have to be accountable. Whether human takeover is feasible depends on the production setting. High-cycle-time assembly lines are difficult, but applications such as feeding materials below the line should generally be workable.
- The commercialization threshold in logistics includes at least 3 nines, or a success rate above 99.9%; the system must also reach the floor for Chinese workers’ cycle times and work with a human-takeover system.
- Scenario selection must satisfy both commercial value and data-generalization value. The scenario should address a real pain point, have broad applicability, and scale from 1,000 products to 10,000 and 100,000, continuously improving model capability. If the system repeatedly handles only 5 or 10 objects in one process, it ultimately approaches nonstandard automation equipment. Technical feasibility is an additional requirement.
- Wang is not convinced by investors who say the embodied industry has detached from fundamentals and is mainly about winning orders. Embodied companies, including Lingxin, remain R&D-driven; if a company can truly deliver a scenario, customers will come. Customers broadly understand that they will eventually enter embodiment, but the number of choices is itself a temptation. Lingxin will therefore narrow its scenario selection this year.
30. Industry Temperature: Waiting for Open Source Is Too Optimistic, Smoke and Mirrors, and the “I’m Very Worried” 1M-Hour Target
- The 2 consensus views in the first half of 2026 are that the brain matters and data matters, but the industry remains uneven. On the claim that domestic brain companies are all waiting for the most advanced U.S. companies to open-source their work, Wang says “everyone is too optimistic.” PI included, no company has fully open-sourced its model or released complete weights; only pieces such as partial datasets have been made public.
- He also relays an unconfirmed industry rumor: the reported 270,000 hours of data “may” contain smoke and mirrors. Domestic brain companies therefore need to recruit top talent and develop their own judgment on technical routes.
- On companies that have raised large sums, claim to be building brains, but are not actually training models, his assessment is “possible,” though he is not certain.
- He partly accepts the skepticism that a demo does not equal model capability. A real, credible demo can still show part of a model’s capability. Specialists can examine operational latency, whether the system is teleoperated, and whether it is controlled entirely by a policy. Some jitter and latency can actually make a demo more credible.
- On the 1M-hour target, Wang says Lingxin was among the earlier domestic companies to propose it, but now everyone is saying it, which makes him uneasy. Do they really have enough insight to judge the target? Lingxin originally back-solved from automotive data that it might need 1B clips, based on roughly 60,000 hours of real data at the time. It still needs to rerun the reasoning from a factual standpoint; the target remains only a hypothesis.
- He believes the industry curve has started running but has not reached the inflection point. By year-end, once the data volume reaches a certain level, model generalization may receive its first-stage validation. An arms race in model iteration based on more data has begun and will last at least 3 years; robots aim to match human performance, but their learning method is completely different from humans’.
- On the view that embodiment has not yet reached GPT-1.0, he says the entry path is different, but the GPT-3.5 moment remains far away. It is too early to say who is the absolute leader; the first elimination round has not even started.
31. Epilogue: Midlife Crisis, OpenAI 2020-2022, and Robots in the Morning Light
- Wang admits he experienced a midlife crisis. Around age 40, what frightened him most was the possibility that what happened in the future would have nothing to do with him. Now that he has joined the embodied wave, he believes there is no endpoint in entrepreneurship at which one is completely liberated; the process itself is the important thing. It echoes the metaphor of “spiraling upward”: needs keep appearing, and the problems keep getting bigger.
- His reading list includes Morgan Housel’s Same as Ever and The Psychology of Money, along with Empire of AI. From the histories of OpenAI and DeepMind, he takes the lesson that founders’ genes shape a company’s path. Demis Hassabis once said the company would win many Nobel Prizes, though no one believed him at the time. DeepMind built AlphaGo but missed large language models.
- If he could align himself with a particular phase of a company, he says neither OpenAI nor DeepMind answered the problem of complex hardware coupling. Tesla has done hardware coupling best, with Apple doing it earlier. In pure models, he would like to be in the period from OpenAI in 2020 through the first half of 2022, moving from GPT-2 toward GPT-3.5—the “most critical stretch, floating beneath the water,” just before intelligence emerged.
- The romantic image is a specific morning. One morning 4-5 months ago, sunlight slanted into the office and illuminated many embodied-device components. What truly moved him and the team was not the grand narrative but each incremental improvement the company made. That is the compounding effect.
- He believes this AI wave will be the most important paradigm shift of roughly the next 50 years. The undercurrent will feel unfamiliar and uncomfortable, but participating in it like surfing may be a good way to move through it.