Pioneers Insight Method Research Author
Vol.194 Industry Watch 36 | Ant Group’s 纪纲 and 李丰: A Heated Debate on Embodied Intelligence and AI Hardware
Back to Episodes

Vol.194 Industry Watch 36 | Ant Group’s 纪纲 and 李丰: A Heated Debate on Embodied Intelligence and AI Hardware

Summary

  • Li Feng’s core argument: this AI cycle is bottlenecked by data, not algorithms. Large language models are built on more than 40 years of internet text, while robots’ motor capabilities rest on 30-40 years of industrial-robot development; China was already the world’s largest industrial-robot market by production and sales in 2013. L2/L3 autonomous driving came from Tesla and China’s new automakers putting cameras and millimeter-wave radar on consumer vehicles. “The biggest gap between where we are now and where we want to be is the infinite amount and infinite dimensions of data beyond text and images”—environment, emotions, vital signs and physical contact. “There is simply no way to skip steps.”
  • Data has only one path to scale: the consumerization of sensors. “Consumers don’t buy sensors for the sake of buying sensors.” Apple and Huawei made cameras ubiquitous, enabling Douyin; GPS ubiquity enabled food delivery; microphone arrays enabled WeChat. The investment implication: the next wave belongs to consumer hardware with new sensors—use demand to drive digitization first, then iterate toward intelligence and personalization. “No device can be AI-native from day one.”
  • Both sides confirmed the embodied-intelligence bubble. Ant Group’s 纪纲 has invested in 8 projects and “may recently be investing in 2, but I feel that’s about the scale”—many companies saw valuations rise 5x in a year with “no observable real progress,” leading him to conclude that “perhaps 80% of companies will be eliminated.” Long term, however, he is categorically bullish: “15 years from now, this will certainly be a larger industry than EVs plus autonomous driving, and every middle-class household in the world will certainly need 1 to 2.”
  • Every humanoid-robot demo is pure locomotion; nobody is demonstrating manipulation. Dancing, backflips, kicking and sports show movement, not the ability to operate objects. Li Feng is skeptical that a large model can serve as the brain or that VLA training can generalize: watching every football match in the world still would not get someone onto a semiprofessional field. “I don’t lack TV data; I lack training data.” Manipulation requires physical models, physical quantities and environmental modeling—and “you don’t have those numbers.”
  • The reason nobody is talking about LLMs and scaling law anymore is that there is no higher-grade, publicly available data of a different type left. 纪纲’s consumer-side observation: OpenAI has 800M weekly active users but average usage is only 17 minutes, suggesting users still see it as “a better search engine”; Dev Day’s Apps SDK, Agent Kit and Codex show commercial intent, betting that in 2 years it becomes an 800M-DAU gateway used 2 hours a day. He also expects OpenAI to work with Luxshare Precision on an extremely cheap data-collection device: “the hardware is free; all it wants is the data.”
  • Tech investing always comes in 3 waves. First comes the technology shift itself, such as LLMs; second, the most imaginative applications—U.S. agents and Chinese robots, with “robots handling the physical world and agents handling the digital world”; third, businesses that can use the technology, prove demand and preferably make money. “The third wave is about to begin, or is already beginning.”
  • DJI and Insta360 are fundamentally the same play. Software-and-algorithm companies use China’s industrial base to bring a high-end product down by half a tier, moving from professional to mass consumer products, much as electronic keyboards once replaced pianos. Both sides agree that glasses are the end-state device, but, as with the iPhone, MP3 players, iPods, BlackBerry and Palm had to come first: “these steps are hard to skip.”
  • 纪纲’s counterargument is worth recording: the moon landing did not wait for every technology to be ready, and teams pursuing AGI or world models could generate hardware spillovers in reverse. Data is an unrecognized mine: “When Baiyun Obo was first mined in 1958, everyone thought it was an iron mine; today we think it is a rare-earth mine.” He is bullish on third-eight-hour data-collection devices and even a “life cheat code”: a Seattle developer collected more than 200G of personal data in 3 days and gave it to Gemini, which suggested he might have cervical spondylosis—something he himself had never noticed.

Deep dive

1. China Has Assembled Four Ingredients—and Is Being Forced to Innovate Without Knowing Where It Is Going

  • Li Feng opened with postwar Japan in the 1970s and 1980s, when the shift from vacuum tubes to transistors prompted the country to convert every mechanical product it could into an electronic one: Shanghai watches became Casio and Seiko; pianos became Yamaha. But Japan lacked sufficient chips and sensors, so it could digitize only so far. “Put it in today’s terms: it was converting fuel cars into new-energy vehicles.” Installing lidar for automated parking was still impossible.
  • Japan was once criticized for overcapacity. But once scale brought unit prices down, penetration multiplied over the next decade: “When our parents got married, a Shanghai watch was still part of a bride’s dowry. By the time I was in middle school, even a teacher’s child could wear a five- or ten-yuan electronic watch.” Exporting capacity was also how Japan built the global market.
  • China today has accumulated 4 advantages Japan did not have at the time: an extremely comprehensive manufacturing chain; a chip-and-sensor supply chain built over the 6-7 years after Huawei and ZTE in 2018; the world’s second-largest goods-consumption market; and the world’s highest circulation efficiency. He redefined “involution” in management terms as “a market that reaches saturation through competition at exceptional speed.”
  • Once these 4 pieces are in place, they “force companies to pursue a great deal of innovation without necessarily knowing the direction.” With imitation, you know whom to copy and only need to make it cheaper. Today, “you no longer know what to do next, but there are many people behind you competing,” so companies must reconnect the opportunities in front of them—some connections work, some do not. China is turning electronic products into digital and intelligent ones, and converting previously nonelectronic products such as guitars into electronic products. Whatever gets pushed beyond China’s borders becomes a world leader—“exactly like Japan back then.”

2. 纪纲’s Three Categories: AI-Native, AI-Carriers and Export Hardware Have Completely Different Valuation Systems

  • 纪纲’s question went straight to the missing part of the framework: supply-chain completeness is a gradual 20-30-year process, so why have smart hardware categories exploded specifically in the past 2 years? He divides what he sees into 3 groups: genuinely AI-native hardware that creates new data increments, such as AI Pin—“there are relatively few truly real examples”; hardware suited to carrying AI, such as glasses and AI-companion products, although the latter can easily collapse into traditional toys; and the largest category, export-oriented consumer hardware, where the competition is over supply chains and distribution.
  • Their valuation logic is worlds apart: AI-driven products are few, but the market prices in potential; export consumer products are valued on operations. 纪纲 asked how Li Feng separates them in investment. Li Feng’s candid answer: “This is what has always tormented me as an investor—you can reason forward to the traits something should have, but you cannot reason forward to what a product with those traits will ultimately become or who will build it.” The answer is to invest in everything that matches at least 50-70% of the traits.

3. Why Nobody Is Talking About LLMs and Scaling Law Anymore

  • Li Feng raised the question: why has nobody discussed changes at the LLM companies over the past 6 months, when 18 months ago everyone was talking about them—and even “scaling law,” a term invented specifically for large models, has disappeared? 纪纲 answered from the consumer side: OpenAI disclosed at Dev Day that it has 800M weekly active users but only 17 minutes of usage time, “only slightly longer than a traditional search engine.” For ordinary users, the novelty has passed; they still see it as a better search engine.
  • But 纪纲 believes Dev Day’s 3-part package showed commercial determination beyond the AGI narrative. Apps SDK plugs solved problems into the framework; Agent Kit lets entrepreneurs tackle unsolved problems within it; and Codex, in his view, is “more for solving long-tail problems—if it can’t solve one, I’ll write it for you on the spot.” If the product becomes 800M DAU with 2 hours of daily use in 2 years, OpenAI will own the user gateway and shared memory across products.
  • From the investor side, 纪纲 attributes the silence to competition entering deep water. DeepSeek’s OCR version, Qwen’s progress and Ant’s recently released trillion-parameter product no longer show their gains on the consumer side: “Whether 峰叔 is twice as smart as me or 100 times as smart as me doesn’t matter anymore, because he is already smarter than me.”
  • Li Feng was blunt in closing the loop: “I didn’t get the answer I wanted.” His real question was whether nobody is talking because “there is no higher-grade, publicly available data of a different type left”—“you can no longer keep progressing along the same path.” He set compute aside for the moment: “Anyway, Nvidia is still seeing its valuation rise.”

4. China and the U.S. Are Hot on Different Things, but Hot Money Has Retreated Again: The Embodied-Intelligence Bubble, Authenticated

  • Li Feng’s cycle review: LLMs were hot 2.5 years ago; embodied-intelligence robots were hot in China 1 year ago, while agents were hot in the U.S. “It was strongly shaped by national characteristics, and correctly so—the U.S. lacks manufacturing, so its excitement leaned software; we have manufacturing and sensors, so ours leaned hardware.” Robots consequently made it into this year’s Two Sessions report. But new financing on both sides has already fallen by a large order of magnitude.
  • 纪纲 supplied live portfolio data: he had previously invested in 8 embodied projects and “may recently be investing in 2, but I feel that’s about the scale.” He believes the bubble is serious: “Many companies saw their valuations rise 5x in a year, but in reality there was no substantive progress.” Li Feng immediately claimed the line: “That is exactly the sentence I came for today.”
  • 纪纲 then added his “but”: this is a cyclical bubble, like the roughly 200 autonomous-driving startups in 2015-16. Several have survived, but remain stuck around L2+; “will it ultimately reach autonomous driving? Absolutely.” His vision is that “15 years from now, this will be a larger industry than EVs plus autonomous driving,” with every middle-class household in the world needing 1 to 2 robots—a highly certain outcome. Otherwise, he would not have told colleagues 2.5 years ago to “invest in as many robots as possible.” But the shakeout will be violent: “Perhaps 80% of companies will be eliminated.”

5. Every Demo Shows Locomotion; Nobody Demonstrates Manipulation

  • 纪纲 exposed the industry’s open secret: review the demos from every prominent company—dancing, backflips, kicks, football, running and sports days—and they are “all pure locomotion.” They are called humanoids and embodied systems, but “none of the other human capabilities, especially manipulation, has been demonstrated.” Li Feng’s judgment was harsher: the robot’s physical movement speed needs to approach human speed, and “none of the technical routes I see today, including the algorithms, can solve this within a few years. I may be too pessimistic.” With no convergence from basic algorithms to data collection and even the robot’s physical architecture, “going straight to industrial maturity is impossible.”
  • His skepticism toward last year’s 2 popular financing narratives—the large model as the robot’s brain, and VLA using broadly available visual data to generalize upper-limb manipulation—focuses on robustness and precision. His analogy: could a loyal fan who has watched every football match in the world and knows every technique and referee rule step onto the field and play at a semiprofessional level? “Obviously not.” Li Feng joked: “If watching were enough, I would have become a badminton champion long ago.”
  • In translation: watching alone may help somewhat, “but it probably cannot get you to real-world manipulation through watching alone,” especially when coordination control, precision, timely response and planning are involved.

6. Three Existing Capabilities, Each Backed by Decades of Data Accumulation

  • Li Feng connected the entire discussion in one line: LLMs exist because “the internet accumulated more than 40 years of text”—“everyone forgot that large models were originally all called large language models.” Robots’ motor capabilities came from 30-40 years of industrial-robot motion control, dual-arm coordination and motor improvements. China was already the world’s largest industrial-robot market by production and sales in 2013, alongside breakthroughs in reinforcement-learning algorithms.
  • Autonomous driving is the same. Fifteen years ago, it was the first huge AI bubble and “everyone said autonomous driving would become widespread very soon.” Ten years ago, people believed L4 was imminent; today, the national standard only permits companies to advertise up to L3. It requires environmental data and a digitized vehicle state—speed, lane position, surrounding vehicles, driver status and payload. “These things probably only became digitizable recently, and we are iterating on that foundation.”
  • His favorite counterexample is 新石器: “A low-key company that delivers lunchboxes in a restricted campus is not the most glamorous, but it raised money in the final wave of every autonomous-driving cycle.” Restricted environments and constrained conditions might make L4 possible; an open environment cannot be solved in one leap. “It was not glamorous. Then it raised RMB4.2B a few days ago, and it was no longer low-key—it moved into the front rank.” The conclusion: “Today’s biggest problem is the same as the large-model problem: it has run out of data.”

7. Consumerizing Sensors: The Only Path from Zero to A

  • The core mechanism of the discussion was this: the data behind all 3 existing capabilities came from “new sensors becoming widespread and reaching consumers, so enough people could help turn them into data.” Text relied on the combination of PCs, keyboards and mice. Driving data came from Tesla putting cameras and high-frequency millimeter-wave radar on consumer cars, and from China’s new automakers entering the race after 2015.
  • The evidence chain is worth preserving in full: “Why did Douyin emerge? Because Apple, Huawei and others made this sensor ubiquitous. Why did food delivery emerge? Because they made GPS ubiquitous. Why did WeChat emerge? Because they made microphone arrays ubiquitous.” The key constraint: “Consumers will not buy a sensor for the sake of buying a sensor.” They buy a product that happens to contain a sensor and naturally converts their needs into data.
  • The gap between today and tomorrow is “far larger than what we originally thought we needed”: human emotions, language, condition, vital signs, environmental modeling, changes in human-environment interaction and an infinite number of physical quantities. He used overhead aerial data as an example: it is difficult for robots to obtain because drones are not yet ubiquitous. “If it became something genuinely flyable that everyone had one of, there would be abundant data from that angle, across dimensions and speeds.”

8. The Moon-Landing Argument: Goal-Driven Progress or Incremental Supply-Chain Completion?

  • 纪纲 announced that he was going to argue the other side: the U.S. did not wait until all the technology was ready or a space station had been built before going to the moon. After the Soviet Union launched a satellite, the U.S. “set a target that was actually unattainable,” and the moon landing drove many technologies in reverse. Likewise, someone may pursue AGI, world models or the final form of embodied intelligence directly, “driving enough industrial development along the way and spilling over many technologies, which then gives smart hardware better foundations today.” The sequence “does not necessarily have to happen in order.”
  • 纪纲 then said there is no simple causal relationship and partly agreed that intermediate links still need to be built, but that does not prevent people from aiming directly at embodied intelligence; the 2 processes interact. He used neural wristbands as an example: once the algorithms exist, the wristband can correlate unstructured electrode data with humans’ limited set of prescribed movements. But because hand-motion data is still missing, devices must first collect more data before a wave of the hand can move a screen.

9. Tech Investing Always Comes in Three Waves; the Third Wave Is Here

  • Li Feng’s cycle model: “Every technology cycle’s investment always comes in these 3 waves. You never skip one, lose one or jump over one.” The first wave invests in the technology shift itself, such as LLMs. The second invests in “the most imaginative thing the technology can do”: “Once large models emerged, agents could replace everything people do on digital devices; everything in the physical world goes to robots, and everything in the digital world goes to agents.” Maximum imagination and maximum difficulty are 2 sides of the same coin—“otherwise it would not be imaginative.”
  • The third wave invests in businesses that “can use the technology, prove demand and preferably make money.” Autonomous driving followed the same path: from the best algorithm teams, to companies connected to automakers, and finally to vehicles already deployed at ports and in industrial parks and generating revenue. “The good news is that the third wave is about to begin, or is already beginning.”
  • That brings back what he calls the “eternal debate”: should investors identify the product that delivers the technology most fully, or start from consumer demand and find a product that is “just a little more AI-enabled than the technology product people used before”?

10. DJI and Insta360 Are Fundamentally the Same Thing: Dropping Half a Tier

  • 纪纲 posed the setup: if you had enough money 10-odd years ago to fund only one company, would you back DJI for creating a new category or Insta360 for improving an existing use case? Li Feng dismantled the distinction: “They are completely the same; DJI just came a few years earlier.” Both were founded by people with computer-algorithm backgrounds. DJI applied flight-control technology and “used China’s manufacturing chain to bring something originally military down by half a tier for professional users.” Insta360 came from image stitching; in 2014 it sold software technology in Nanjing during the first VR wave, when Facebook had just bought Oculus, but could not make money. It moved to Shenzhen and grafted stitching technology onto the manufacturing chain to build 360-degree action cameras, initially selling to Western consumers educated by GoPro.
  • The formula is to use industrial-chain capabilities to bring a high-end product down by half a tier, then, once a profitable flywheel is established, lower it another half-tier into the semiprofessional market, and now another half-tier toward mass consumption. The reference point is still the electronic keyboard: “In the 1980s, what kind of family could afford piano lessons for a child? With an electronic keyboard, a middle-class family could do it in the 1990s.”

11. Three Eight-Hour Blocks and the “Life Cheat Code”: 纪纲 Bets on Incremental Data Collection

  • 纪纲’s internal framework is the “3 eight-hour blocks.” Sleep produces large amounts of data but offers no interaction. Screen time is already occupied: “If AI wants to build a better ride-hailing or food-delivery app today, the opportunity is small. The previous generation was not AI-driven, but algorithms already solved those problems extremely well.” The real increment is the third 8-hour block—for example, a conversation between 2 people, which has traditionally been very difficult to capture.
  • He imagines a progression from Pro-style audio-recording products to visual analysis of facial expressions—“What does it mean that you are biting your little finger? Are you thinking?”—and eventually what he calls a “life cheat code.” In an interview, it could read the interviewer’s expression and tell you how to answer the next question. He acknowledges that “the final boss should be glasses,” but battery life and compute have not been solved, and capture, display and interaction cannot all be packed into one device today. Hence compromises such as a chest-mounted “Lupi” device: slightly worse perspective and image quality in exchange for battery life and the initial data-collection function.

12. Glasses Are the End State, but Not Yet: Digitize First, Then Add Intelligence and Personalization

  • Li Feng agrees that “glasses will definitely be one of the ultimate major devices,” but maps the intermediate products into a matrix: either use sensors to collect multidimensional human and environmental data, or use cameras to capture multiple scenes, states and emotions. Ideally, the endpoint can tolerate a chip for edge-cloud integration, despite the power, size and cost challenges. Then the product should “iterate from digitization to intelligence and then personalization.” His methodological summary: “Do not define demand from the data layer; define the product from demand and obtain the data.” Consumers “will absolutely not” buy a device defined solely to sell data to robot companies.
  • The historical evidence against skipping steps is the iPhone lineage. Before Apple’s phone came the iPod; between the iPod and Apple came BlackBerry and Palm; before the iPod came “China’s educational MP3s”; before that came indestructible Nokia phones. “It was not until the third generation that you began to think the iPhone was a good phone.” Even if you are Steve Jobs, “it is hard to educate consumers by skipping directly ahead.” Cameras followed the same path: Nokia and Motorola had rear cameras, but nobody used them. The iPhone first made 2-3 megapixels useful, then locked users in by storing photos in the cloud. “If you try to build an extreme camera from day one, that may be challenging.”
  • This is why he expects OpenAI to move into hardware: “OpenAI will eventually find Luxshare Precision to make hardware; it will eventually make some kind of plot, as the source puts it.” For OpenAI, vast amounts of text with context and scenes would be extremely valuable. “I suspect it will ultimately define something so cheap that the hardware is free; all it wants is the data.”

13. The Final Argument: Data Is Rare-Earth Ore We Have Not Yet Recognized

  • Before wrapping up, 纪纲 made one last argument “to ensure we have something to talk about next time.” Sleep, an intensely subjective experience, has been quantified, albeit inaccurately: “It gives you a score of 80 and you feel you slept terribly, but the emotional comfort is real. I have never reached 80; my highest was 78.” Yet it still creates value. This is the process of making data visible.
  • His most important example is a solo developer in Seattle. He first wore Insta360’s “Ultra Go” and found the battery life inadequate, then hand-built a microcontroller with a camera that took 1 photo a minute and 15-second video every 3 minutes. He set an iWatch to record audio 24 hours a day and collected more than 200G of data in 3 days, then fed it to a version of Gemini from a few months earlier. The most surprising result was that the model told him, “You probably have cervical spondylosis, because I can see that your posture is bad in many situations”—something he himself had never noticed.
  • The closing metaphor was the mine: “When Baiyun Obo was first mined in 1958, everyone thought it was an iron mine; today, we think Baiyun Obo is a rare-earth mine.” Until uncollected data is captured, we do not understand what value it may create. Once humans are fully embodied, “perhaps every second of your data will be valuable.” Li Feng did not pick up the argument.