Huawei 'Genius Youngster' Enters the Embodied AI Marathon
Huawei 'Genius Youngster' Enters the Embodied AI Marathon
Summary
- 黄青虬 was likely the last of Huawei’s 2020 “Genius Youngsters” to enter embodied AI, because he had an “unfinished mission”: taking Huawei’s single-stage end-to-end intelligent-driving system from research to commercial mass production in 2025. He sees GPT as solving only virtual-world problems: LLMs operate through turn-based dialogue, where the environment does not change while answering, whereas embodied systems and vehicles are real-time, closed-loop systems (“If you drive fast, the car next to you won’t dare cut in; if you drive slowly, someone else will cut in”). He waited until 2 milestones emerged in mid-to-late 2025—the evolution of LLMs from GPT to coding and agents, and the commercial mass-production validation of single-stage end-to-end intelligent driving—before founding 墨甲智能 at the end of 2025.
- Embodied AI is a marathon, not a 100-meter dash, which explains why new teams can still emerge after 3 years of a crowded race. “Being a few lengths behind at the start doesn’t matter; only in a 100-meter sprint do you need to align the starting line precisely.” The technical paradigm has also not converged: data collection is shifting from teleoperation to body-free capture, raising aggregate collection efficiency by more than 10x; “what everyone accumulated over 3 years in the past may be caught up in 3 months.”
- The company closed more than RMB1B in angel funding within 6 months, from investors including Alibaba, Tencent, BlueRun, Legend Capital, Gaorong, and Source Code, and its compute spending should already be in the RMB10M-per-month range—possibly at least top 5 in the industry. His funding yardstick for training the “brain” is straightforward: 1,000 GPUs is the current threshold; doing a reasonably good job may require 10,000 GPUs to have any chance of producing a breakthrough. At roughly RMB10M/month for 1,000 GPUs, or more than RMB100M a year, a company needs to raise at least RMB300M to support 2 years of compute; without that amount, “you’re basically not in the business of training the ‘brain.’”
- His non-consensus view on data is that quality matters far more than quantity; the industry data he has seen is “not high quality and highly uneven.” His 3-real rule is “a real person doing real work in a real setting” through in-the-wild collection. He criticizes professional data collectors for staging tasks—“wiping a table to produce the wiping motion, not to get the table clean”—because “the data is just a staged trajectory, so the model naturally learns a staged trajectory.”
- The commercial path runs from enterprise as a springboard to consumers as the end state, with hotel or serviced-apartment room cleaning a possible first stop. The setting resembles a home and contains more than 160 subtasks; completing 30%-50% would be enough to enter the field and feed data back. He is not developing components in-house for now: from his own test, “I can do 70% of the work in a hotel wearing a gripper, while the algorithm may not even manage 5% today”—the bottleneck is the algorithm, not the actuator.
- On Unitree, he says reinforcement learning has lowered the barrier to motion-control algorithms, while its bigger lead comes from hardware consistency and parameter identification that have made the sim-to-real gap exceptionally small—and that lead “can be caught” because it is mostly engineering experience, not highly uncertain technology. The real uncertainty is in the brain; embodied AI’s scaling law is also harder than the one for LLMs: edge compute limits parameter counts, while each frame carries more multimodal information.
- His sharpest pushback is against “the first year of mass production” and PMF narratives: occupying a vertical use case “has not created a market monopoly or moat,” while hardware that can run 7×24 for a year without breaking is something he believes no robotics company can do today. “Overclaiming has very little meaning. If something cannot run consistently and reliably 7×24, cannot be produced at 10,000 units with every unit consistent, then in my view it is a prototype.”
Deep dive
1. The Last “Genius Youngster” to Jump In: GPT Alone Cannot Deliver Embodied AI
- Among Huawei’s 2020 “Genius Youngsters,” 彭志辉 at 智元 and 丁文超 at 踏实 had already left to start companies. 丁文超 once worked with 黄青虬 in the same conference room, “happening to piece together the 2 most important parts of robotics”—perception and planning/control. 黄青虬 should be the last to make the move so far: “I had an unfinished mission at Huawei”—taking end-to-end intelligent driving from 2023 research to commercial mass production in 2025, a process that took 1.5 years.
- His sober take on the ChatGPT moment was that GPT solves virtual-world problems, with 2 fundamental differences from embodied systems. First, LLMs operate through turn-based dialogue—“being a few seconds late with an answer is fine”—whereas vehicles and robots are real-time systems making decisions every second. Second, embodied systems are closed-loop: every decision changes the environment, while the environment remains unchanged when an LLM produces an answer.
- His conclusion: “There still needs to be an AI that truly solves the closed-loop system and validates the full chain that drives low-level control before embodied AI can really happen.” That is why when 彭志辉 started his company in 2023, 黄青虬 “never even considered” following him—the prerequisite technology was not ready.
2. 2 Inflection Points Arrived in Mid-to-Late 2025
- The first was the evolution of LLMs from GPT to coding and agents. The second was single-stage end-to-end intelligent driving: “input an image, output a control signal.” Using AI to do this proved workable, generalizable, and scalable for mass commercialization. Together, the 2 developments meant that both high-level semantic understanding and low-level control could be driven by a single AI system.
- That led to the founding of 墨甲智能 at the end of 2025. “墨” comes from Mozi, who in the Warring States period built all kinds of engineering devices, as well as equipment for warfare and transport. “甲” refers to the starting point and an explosive beginning—the Big Bang. His own name also has a source: “Green dragons and purple swallows sit in the spring breeze” (「青虬紫燕坐春风」). 虬 means dragon; his father chose it in the hope that his son would become a dragon, but without being too showy. “青” also evokes the idea of surpassing one’s teacher.
3. More Than RMB1B in Angel Funding in 6 Months; Compute Spending May Rank Top 5
- After news of his departure broke, “my WeChat was flooded with new contacts.” The company closed more than RMB1B in angel funding in 6 months, from Alibaba, Tencent, BlueRun, Legend Capital, Gaorong, 中科创星, and Source Code Capital. The level of excitement was “completely beyond expectations,” and the company can currently “pick its investors.” Patience is the first screening criterion, because this is a “huge market and an exceptionally long race.”
- The biggest uses of capital are talent and compute. Compute spending should already be in the RMB10M-per-month range. “We’ve only been around for 6 months, but our compute investment may already rank at least No. 5 in the industry”—which also suggests that “the industry’s algorithm spending has not kept pace with the sector’s level of excitement.”
- He gives his confidence in the startup at least 90 points. “Investors seem never to have asked me this question. They may simply assume I can make it.”
4. Starting a Company Is Not Hard; Succeeding Is
- The host asked directly whether the flood of embodied-AI startups meant the barriers were low. His answer: “The barrier to starting a company is not high, but the barrier to building a successful company is very high.” Embodied AI is a complex hardware-software system that demands exceptionally strong vertical-integration capabilities.
- He agrees with the logic behind dollar investors’ preference for founders from intelligent driving. After 6 months in the field, he found the embodied-AI stack “particularly, particularly similar” to autonomous driving: a low-latency, high-dynamics integrated hardware-software system; an image-to-action algorithmic mapping; and a data loop spanning collection, cleaning, automated labeling, training, and evaluation. The 3 layers of methodology are reusable.
5. The Marathon Thesis: Why New Teams Are Still Appearing in 2025 and 2026
- The core metaphor is simple: “Embodied AI is a marathon. Being a few lengths behind at the start doesn’t matter; only in a 100-meter sprint do you need to align the starting line precisely.” What decides the outcome is organizational cohesion and judgment—and taste—at each paradigm shift.
- More importantly, the paradigm has not converged. Data collection is moving from teleoperation to body-free capture, with aggregate acquisition efficiency improving by more than 10x: “what everyone accumulated over 3 years in the past may be caught up in 3 months.” There is no fixed market structure and no exceptionally high barrier to entry.
- The convergence of team profiles reflects how fiercely contested the race has become. 3 or 4 years ago, a team needed only commercialization and scientific credentials; today it also needs engineering, industrialization, and productization. The result is an all-around, “bucket-shaped” team.
6. The Industry’s Pathology: Inflated Egos and 2 or 3 Companies Splitting From Every Leader
- The most visible problem, he says, is talent fragmentation: leading companies sometimes split into 2 or even 3 companies after operating for a while. The excitement gives everyone license to fantasize: “He’s not that different from me. He can raise money, so can I?” The flattering description is optimism and a willingness to take risks; the less flattering one is excessive restlessness.
- His answer is cultural selection: “humble but confident”—confident in technical judgments while open to other views and able to find common ground amid differences; and “gentle but firm”—holding to sound positions while resolving conflict gently. The identification mechanism is simple: “people with the same temperament attract each other,” provided the observation period is long enough.
- A personal footnote: he served as class monitor continuously from primary school through high school. “Mischievous kids would challenge me, but when they wanted to copy homework, they would give in.”
7. No Need to Overtake on the Inside; Just Run 1 Centimeter Farther Every Step
- Facing competitors with a 2- or 3-year head start, he says: “The finish line is more than 100 kilometers away. I only need to run 1 centimeter farther than them with each step. In 6 months, 1 year, or 5 years, I’ll gradually move ahead.”
- What he demands from the team is dirty work, hard work, and unglamorous craftsmanship: squeeze reliability out of every hardware detail, and scrutinize every small module in the algorithms and data. “Verify precisely that both accuracy and efficiency are sound. Once the foundation is solid, progress will come quickly.”
8. Consumers Are the Destination, Enterprise the Springboard: Why Hotel Cleaning May Come First
- The consumer logic is that any general-purpose category ultimately has to reach consumers to access a sufficiently large market and profit pool, and the most universal setting is the home. But today, robots cannot even fold clothes neatly. Entering the consumer market prematurely would damage the brand; privacy also makes data feedback difficult, while “AI definitely needs data feedback to keep iterating.” The plan is therefore to use enterprise deployments as a springboard before eventually reaching consumers.
- Generalizability is the first criterion for choosing a setting: “The setting you choose determines what data you collect, which determines what kind of base model you can train.” That is why hotel or serviced-apartment room cleaning is a possible choice: it resembles a household, is sufficiently varied and complex, and generates high-value feedback data.
- The second criterion is the ability to enter in stages. Room cleaning is a collection of more than 160 subtasks. “Even if all it can do is pick up trash from the floor or replace the bottled water in the refrigerator,” completing 30% or 50% is enough to enter a real setting, feed data back, and expose real problems.
9. The Data War Has Begun, but He Is Betting on Quality, Not Quantity
- The host observes that the data war began in 2026: large tech companies and model companies are “buying data at any cost,” while new data-collection firms are touting 10M hours or 100M hours. His non-consensus view is that “data quality matters far more than data quantity,” and the embodied-AI data he has seen is “highly uneven,” with no common standard guaranteeing that a particular collection method will train a strong model.
- He maps the evolution of collection paradigms as follows: teleoperation with an embodied robot—synchronous VR control, low efficiency, high cost, and poor diversity in staged indoor settings—followed by body-free capture using a UMI gripper, motion-capture or EMG gloves, and head-mounted cameras recording first-person video. The main problems being solved are collection efficiency and generalization across settings. The company uses vendor sample data, purchases a small amount of data, and builds its own devices for collection.
- His 3-real rule is “a real person doing real work in a real setting”—in-the-wild collection. He criticizes professional collectors who stage tasks: “wiping a table to produce the wiping motion, not to get the table clean.” The resulting trajectories are distorted. “Data quality determines what kind of intelligence the model can learn.”
10. Embodied AI Has a Scaling Law, but It Is Harder Than the LLM Version
- He believes a scaling law exists because embodied AI follows the same AI training paradigm: provide the input and output, and hand everything in between to the model. The paradigm has already been validated in CV, or AI 1.0; LLMs, or AI 2.0; and autonomous driving.
- The challenge is a contradiction. Embodied systems run at the edge, where compute is constrained and parameter counts cannot reach hundreds of billions or thousands of billions. Yet they continuously ingest high-definition video, speech, text, and other multimodal inputs in real time. “The amount of data absorbed in each frame is greater than in a large language model.” Fewer parameters but more information creates a bigger challenge for model architecture.
11. “Copying Homework Doesn’t Get You a Good Grade”: The Algorithm and Compute Thresholds
- He rejects the view that algorithms are unimportant, that money can simply buy talent, and that throwing enough resources at the problem will produce a breakthrough. The algorithm architecture itself is easy to copy, “but the algorithms, talent, and organizational atmosphere accumulated while exploring the algorithms cannot be copied.” Expecting to poach a few people at the moment of breakthrough and immediately fuse them into a working system “doesn’t really hold.” The host picked up the metaphor: “The class monitor says copying homework won’t get you a good grade.”
- His compute benchmark for training the brain: 1,000 GPUs is the current threshold; doing a reasonably good job may require 10,000 GPUs to have any chance of a breakthrough. At roughly RMB10M/month for 1,000 GPUs, or more than RMB100M a year, a company needs to raise at least RMB300M to fund 2 years of compute. That is also the early-stage financing line separating the leaders. If a company cannot raise that amount, he confirms it is “basically not in the business of training the ‘brain.’”
12. Build Full-Stack, but Do Not Chase a High In-House Ratio: The Bottleneck Is the Algorithm, Not the Actuator
- The commercial case for full-stack integration is that software licensing alone is currently unlikely to generate adequate returns. More often, the product needs integrated hardware and software to capture greater commercial value. But in-house development has a clear boundary: “We will only develop components ourselves when the maturity or rate of improvement of component suppliers falls behind us.”
- The company is not developing components in-house for now. His evidence is a firsthand test: “I can do 70% of the work in a hotel wearing a gripper, but with the algorithm today, it may not even manage 5%.” The bottleneck is not the actuator. “I’ll first take the algorithm from 5% to 65%.”
13. Do Not Get Swept Up in the “First Year of Mass Production”: 7×24 Operation, No Breakdowns for a Year, and Consistent Production
- He does not want to be driven by competitors’ claims of mass production or a first PMF. Even if there are enough vertical settings to occupy a position, “taking one spot has not created a market monopoly or moat.”
- The 2 foundational capabilities he is most eager to solve are hardware reliability—“Can it run 7×24 without breaking for an entire year? I believe no robotics company can do that today”—and base-model generalization. If every workstation requires large amounts of labor, compute, and data for fine-tuning, “it is not scalable up from the setting itself.”
- His sharpest verdict came at the start of the episode: “Overclaiming has very little meaning. If something cannot run consistently and reliably 7×24, cannot be produced at 10,000 units with every unit consistent, then in my view it is a prototype.”
14. The Vulcan Team Legacy: Do Not Follow the Crowd; Details Decide the Outcome
- A Tsinghua Automation graduate and alumnus of 赵明国’s 508 Lab Vulcan team, he spent 60% of his undergraduate years in the lab. In his third year, he took over the understaffed perception function and became team captain. 赵明国’s unconventional approach left the deepest mark: while the industry favored ZMP static gaits, represented by ASIMO, 赵坚持 passive walking and virtual-slope dynamic gaits, which save energy and move faster. 黄青虬 still insists on incorporating kinematics and dynamics into AI-based motion control. His credo: “Don’t follow the crowd.”
- The most important lesson he took away was that details decide outcomes. “A gait that cannot be tuned properly may not be an algorithm problem; it may simply be that a screw is not tightened.” Complex systems must be broken down into the smallest modules, with each input and output tested individually. Because hardware, software, and algorithms are connected in series, errors are amplified.
- Asked whether the era’s perception-decision-control split resembled today’s VLA, he explained that it was mostly V: L was replaced by rules, while A was action control. It was a state machine: “See the ball and the goal, shoot; see the ball but not the goal, turn around once.”
15. On Unitree: RL Lowered the Motion-Control Barrier; the Real Lead Is Sim-to-Real
- Motion control has evolved from first-generation ZMP to second-generation MPC+WBC, represented by Boston Dynamics’ Atlas, and then to today’s fourth generation based on reinforcement learning. “The barrier is relatively lower now.” Engineers no longer need to finely model a high-dimensional optimization problem mathematically; they set a reward mechanism and let the model repeatedly roll out in simulation.
- Unitree “has genuinely done an excellent job,” but its bigger lead is the unusually small sim-to-real gap. The foundation is hardware: consistency such that “every one of 100 robots produced has the same parameters,” combined with parameter identification fed into the simulator. Can that lead be caught? “Yes. It is more about engineering experience; there is no particularly powerful uncertainty technology in it.” Unitree, of course, is also moving forward.
- The real uncertainty is in the brain. The source of his 90-point confidence is: “Although I don’t know what it will ultimately look like, if we build a solid foundation, produce high-quality data, and have a good enough team, we can figure it out through trial and error.”
16. Tang Xiao’ou and MM Lab: Black-Sheep Culture and Optimism as a Method
- He interned at SenseTime during his senior year, when the company was just getting started, in a room at Wenjin International Apartments: “the living room was the office, the bedroom was where we slept, and more than 10 people were working on face recognition and face detection.” That was his entry into CV. “I may be face-blind, yet it could tell people apart.” The host describes MM Lab as SenseTime’s “technical mother ship.”
- Tang Xiao’ou’s black-sheep culture used a black sheep as its mascot: “everyone could be an individual,” in opposition to an assembly-line paper-production factory that creates homogeneous research. He admired 2 types of work: research that closes a field by producing a definitive algorithm, and research that opens a new field by posing a valuable question no one had thought of before.
- 黄青虬 took away 2 principles: pursue excellence—“aim high, and you’ll get somewhere in the middle”—and optimism. Optimism “should be a method,” not an innate trait; it can be shaped through experience. “With the right mindset, you can always get through it.” Tang’s death was “a tremendous loss” to the community. In his view, Tang Xiao’ou played a pivotal role, at minimum in the Chinese community and in China’s AI industry.