Pioneers Insight Method Research Author
166: 许华哲 Returns to Embodied-AI Entrepreneurship: He Doesn't Want to Miss the Biggest Watermelon
Back to Episodes

166: 许华哲 Returns to Embodied-AI Entrepreneurship: He Doesn't Want to Miss the Biggest Watermelon

Summary

  • 许华哲 believes scientist-founded companies have gone mainstream, and that technological conviction and long-term vision matter most in the AI era. Citing Ilya, Hassabis, Hinton, LeCun, OpenAI and DeepMind, he argues that the judgments that truly change the world often come from staying committed to long-term goals; but 破壳 is not trying to become a large company’s research arm—it wants to build an independent, consumer-facing robotics brand for the long haul.
  • 破壳 wants to define the home-robot category. Apple represents defining a new category; Xiaomi represents involving users and finding the broadest common ground. The first phase will involve its own hardware, AI model and product definition, likely starting with a wheeled dual-arm robot, with a legged platform potentially developed in parallel, targeting roughly 2 hours of battery life and homes worldwide.
  • His AI-native approach rejects solving traditional robotics one task at a time, closing the data loop in constrained environments, and assembling small models. AI is fundamentally an inductive engine; data type and diversity matter more than homogeneous volume collected in a closed environment. 破壳 wants a unified model, scaled post-training and reinforcement learning to let robots explore, evaluate and use data while preserving multi-task generalization.
  • Early products will control risk by drawing clear boundaries. Wheeled robots cannot truly climb stairs, while weight, batteries, carrying loads, docking and overseas deployment remain unresolved. The first generation will not provide services involving direct contact with human bodies; high-risk tasks such as feeding, holding babies, turning people over and massage will come much later.
  • 许华哲 believes in ecosystems, concentrating resources on AI and product rather than building every component in-house. Based on the current data volume, model pretraining and post-training could require RMB100M–RMB200M a year; as the data scales, the budget could approach that of a large model. He hopes real home robots will already be in users’ homes roughly 2 years from now, around early 2028.
  • China has not missed the window for general intelligence, he says, but the industry must not let near-term shipments and demonstrations dictate its direction. He warns against selling data to US competitors, blindly mass-producing robots without disclosing daily active usage, and endlessly optimizing dance routines. He is watching PI, Generalist, Sunday and Figure most closely on intelligence, and roughly estimates that fully human-level general intelligence may still be about 5 years away.
  • Robots are not a hobby for him but a form of love and a mission. He chose robotics because it is harder and potentially more consequential than game AI; the odds of startup success may be very low, but as long as he gives everything he has, creates something valuable and uses technology to help more people, he is willing to pay “all the costs” (所有的代价).

Deep dive

9. Scientist Founders and the Consumer-Brand Goal

  • 许华哲 says the market worried several years ago that scientist-founders cared about technology but not business, and that technical talent could always return to academia or take an engineering job rather than go all in. On AI’s time scale, he considers that view “ancient history.”
  • Citing Ilya, Hassabis, Hinton and LeCun, he argues that the AI era especially needs scientists’ breadth of vision. OpenAI’s early commitment to scaling data was not widely understood until GPT, and especially ChatGPT, arrived; in his view, the judgments and convictions of scientists are what make world-changing breakthroughs possible.
  • DeepMind started around 2010, arguably too early, and simply surviving before ImageNet emerged was an achievement. After Google acquired it, the company retained relatively independent operations and became an important part of Google’s AI effort.
  • But 许华哲 does not see himself as purely a scientist, and acquisition by a large company is not 破壳’s objective. He would rather build an enduring, consumer-facing company with a brand of its own.

10. 破壳 Wants an Apple-Style Category, with Xiaomi-Style Participation

  • Apple was 许华哲’s first reference point: the iPhone defined the smartphone as a new category. In his view, phones converged after the iPhone into mostly black slabs with large screens, while home robots still lack a common product definition. He wants that definition to emerge from 破壳.
  • Xiaomi offers a different model. In 《参与感》, 许华哲 saw what an ideal consumer company could look like: close to users in its early days, attentive to users and employees, and willing to involve users in product design. Robots will shape future lifestyles, so one company should not define them entirely behind closed doors; the goal should be to find the “greatest common denominator of humanity.”
  • 程曼祺 brought up 乔布斯’s idea of designing with one’s back to the user. 许华哲’s view is that before a product exists, user interviews tend to produce mutually conflicting requests about height, length and other specifics, so a company must first provide a framework. But if the product proves broadly unusable in homes, it cannot cling to its original theory.

11. Never Having Had a Job Brings Freedom—and a Standardization Gap

  • 许华哲 has never worked in a conventional industrial company. His experience consists mainly of research and teaching at Tsinghua’s Institute for Interdisciplinary Information Sciences and more than 2 years as an entrepreneur at 星海图. He acknowledges gaps in his familiarity with standardized processes, but believes they can be learned and absorbed.
  • The benefit of never having had a regular job, he says, is freedom from inherited constraints: workflows, organizational structures and evaluation systems can all be reimagined. He gives the example of AI-assisted writing: what once might have been labeled plagiarism could now merit a perfect score if a student had not handwritten a single line of code but delivered work that was 80 out of 100 in completeness.
  • 破壳’s new office is intended to feel like a home. The living room can host meetings, robot training and podcast recording, while the kitchen and other spaces should reproduce real domestic settings. The team remains tiny, more like a “robotics interest group” made up of software and hardware engineers, lab members and students building the first product.
  • The team will expand, but 许华哲 wants people who can take work in their respective fields to the extreme. Priority hires include hardware engineers with consumer-product experience, top-tier AI researchers, and people capable of helping define products and future ways of living.

12. Trust the Ecosystem, and Focus the Company on AI and Product

  • One important point of agreement between 许华哲 and Eric Zhang is that a robotics company does not need to do everything internally. Camera frame rates and the winding of motor coils, for example, are not necessarily problems the company needs to solve itself.
  • He puts “extremeness”—doing things to the highest standard—at the top of the company’s culture because individuals and organizations have limited attention. When attention is concentrated on A, B will be relatively mediocre; a company that wants to be exceptional cannot try to cover everything.
  • 破壳’s 2 priorities are AI and product: intelligence, and whether the robot body combined with AI can genuinely serve users and keep them coming back. It can rely on the ecosystem for components such as world models and VLMs rather than rebuilding everything from scratch.

13. The First Layer of AI-Native Means “Not Robotics”

  • 许华哲 rejects the traditional robotics practice of solving extremely difficult and impressive problems one at a time when each maps to only a single task. Such systems may use rule-based methods or small deep-learning models trained for one task, but the tasks are enumerated in advance and rarely generalize to the next one.
  • His examples include somersaults, martial arts, balancing 3 spinning tops on a stick, multiple robotic arms pouring drinks at high speed, and a band made up of Kuka arms. These demonstrations may be genuinely impressive and may require sophisticated mathematics and control, but they are “definitely not the path to general-purpose Physical AGI.”
  • His point is not to deny the difficulty of the engineering. An arm programmed to play bass can be highly capable, but it still solves a problem that was defined in advance.

14. The Second Layer of AI-Native Means “Not Autonomous Driving”

  • A common autonomous-driving approach is to close the data loop in a local setting: collect enough data at Wudaokou first, then expand to Zhongguancun and Haidian. 许华哲 worries that embodied intelligence could follow the same pattern—solve one small task, then use that task’s data to solve the next one.
  • His core view is that “AI is fundamentally an inductive engine.” If it only sees people using cups, it may infer that everyone needs a cup; only after seeing babies drink milk without ever using one will it understand that people may or may not use cups.
  • That is why data type and diversity matter so much. Once a model forms an overly strong induction, adding counterexamples later can be extremely painful; the better approach is to expose it to diverse cases together from the start.

15. The Third Layer of AI-Native Rejects “Caveman Deep Learning”

  • “Not caveman deep learning” means not expecting one small model stacked on another, 100 times over, to become the equivalent of a large model. A small-model stack may work if the task is moving the same kind of brick every day; starting from that architecture is the wrong path if the goal is general intelligence.
  • 许华哲 believes part of the industry’s disagreement comes from people who do not truly believe Physical AGI will emerge and see humanoid robots as mechanical arms with a human shape. Others believe in scaling laws but reduce scaling to collecting more homogeneous data in a fixed, closed environment.
  • He also mentions a friend who worked on traditional robotics optimization and once dismissed deep learning as “alchemy.” After using what the transcript calls “open cloud” or “cloud code,” the friend began to reconsider, saying that if even an alchemist could achieve results of that caliber, the approach was already powerful enough. The specific tool name was not clear in the conversation.

16. Jagged Intelligence Requires Both Experience and Product Boundaries

  • 许华哲 acknowledges that “jagged intelligence”—frontier models performing brilliantly on hard problems but failing on simple ones—will be a problem, but believes experience must ultimately settle it. He cites 何恺明’s analogy of an experienced driver: even a veteran cannot guarantee never crashing, but years without a major accident create trust.
  • Data-driven embodied models will have holes, and those gaps can only be filled gradually in operation. Product design must simultaneously make clear what the robot will not do, preventing a basic error from becoming a major accident.
  • 破壳 will not initially offer services involving direct contact with human bodies, including wiping an elderly person’s body, turning them over or lifting them, holding babies, or massage. Feeding may address a real need, but its requirements for accuracy and force control, along with policy and user-psychology concerns, mean it will be deferred to a much later stage.
  • He compares product boundaries to supervising a child. Children make basic mistakes too, but if they are kept away from matches, gas canisters and sharp objects, the worst outcome may be a broken remote control.

17. The First Product Starts with Wheeled Dual Arms and Targets Homes Worldwide

  • 破壳 still wants to build a general-purpose robot with a humanoid form, but the first phase may start with a wheeled dual-arm platform; a legged version may also be developed in parallel. Wheels have practical advantages, but a home with stairs could require 1 wheeled robot per floor, since it cannot truly climb between them.
  • Battery life has not been finalized. 许华哲 currently imagines a minimum of roughly 2 hours. Homes rarely require 2 continuous hours of work, and the robot can return to recharge before resuming its tasks.
  • The target market is households worldwide. He wants a home model to behave like a multilingual model: acquire capabilities from one data distribution, then adapt to a new environment with relatively little additional data from another language or household.
  • The robot body still has to handle the particulars of real homes: weight reduction, whether it can carry objects, whether its battery can enter a home, restrictions on large batteries when going overseas, the robot’s overall weight, and whether a user can return it to the dock after a shutdown.

18. Reinforcement Learning’s Value Lies in Exploring, Evaluating and Using Data

  • 破壳’s model strategy can be summarized in 3 steps: give the robot the ability to explore, interact with the world and generate data; have it evaluate that data; then find ways to use the data effectively.
  • 许华哲 believes the industry may be underestimating reinforcement learning, especially the evaluation step. Data can be high-quality, suboptimal or failed; if all of it is fed into training indiscriminately, a high share of bad data can degrade the policy.
  • Suboptimal data is not necessarily useless, and failed data may also be exploitable. The robot should interact with the world itself and judge different data and behaviors—a potential point of divergence between 破壳 and other approaches.

19. The Action Model Must Be Unified, and Post-Training Must Preserve Generalization

  • The top-level VLM in a software system may be layered, but 许华哲 insists that the parts directly responsible for action, behavior and getting work done should use a unified model. Training on different data together is what can produce generalization across tasks, rather than adding one task-specific model after another indefinitely.
  • He observes that the industry’s pretraining stage is usually a unified, end-to-end large model, while post-training often ends after optimizing 1 task. The broad capabilities of the pretrained model then contract into a narrow task after post-training.
  • 破壳 wants to scale post-training while improving efficiency and success rates without sacrificing generalization. Early on, every task may be mediocre, but he would rather accept “shared mediocrity” and see all tasks improve gradually than have 1 exceptional skill and no ability in the rest.
  • When he said “1 more year,” he explicitly noted that it was just an offhand example; the actual timeline is difficult to predict.

20. A General-Purpose Approach Needs Patient Capital—and Visible Progress

  • Once the company is operating, the team, investors and market will all want step-by-step evidence of progress. Technically, it may be rational for multiple tasks to improve together; organizationally and in terms of resource allocation, that is much harder to explain.
  • 许华哲 calls this “mutual filtering.” When OpenAI insisted on scaling and piling on data in its early days, most people may not have believed the approach, but some who did were willing to come along. 破壳 will look for those partners, while still showing progress rather than remaining completely opaque; intermediate milestones are already planned.
  • A group of investors told him that they may care less about whether the company can produce a short-term result in 3 or 6 months than whether the final product is “the biggest possible thing.” 破壳 will nevertheless show progress and intermediate results.
  • The target cadence is to give the robot some capabilities in a home-like office after roughly 1 year, and to have users operating robots in their homes after roughly 2 years. 程曼祺 has agreed to return then and ask whether the company delivered.

21. The Current Model Budget Is RMB100M–RMB200M a Year

  • Based on the current data volume, 许华哲 considers annual spending of RMB100M–RMB200M on embodied-AI pretraining and post-training reasonable. The estimate depends on the amount of data and is not a fixed total cost.
  • As the data volume grows, he expects the spending scale to approach that of large models. Once the path begins to converge, the scale of resources may become more important.
  • Larger companies currently cannot easily step away from their core businesses; they can only establish small outposts or labs. Once home robots are proven capable of becoming a core business, those companies will enter the field.

22. Data and Models Will Get Larger; the Model Path and Robot Body Are Still Uncertain

  • 许华哲 sees several relatively certain future changes: data will increasingly shift toward video, grow in volume and become more evenly distributed; models will grow with the data; and users will gradually accept robots as life terminals akin to smartphones.
  • The real uncertainties are what kind of model can absorb the data and what kind of robot body users will accept. Hardware is a bottleneck, but a solvable one: if 1 motor is missing, it can most likely be built or sourced, given enough time and engineering effort.
  • Robots need more than semantic priors; they also need physical priors, such as whether an object will bounce or simply fall after being dropped. Robot pretraining can build action and interaction priors, while post-training teaches it how to perform tasks better.
  • A world model can serve as a backbone, generator or data generator; it can also predict the next frame and then solve the inverse problem to obtain an action. 许华哲 says this remains under exploration. 破壳 currently leans toward using a world model as the backbone, and sees no need to build a world model or VLM from scratch.

23. He Watches Global Peers for Intelligence—and Remains Skeptical of Figure

  • 许华哲 is watching the intelligence progress of PI, Generalist, Sunday and Figure. In his view, the DNA of the other 3 companies determines, in some sense, that what they build will be the real thing, and their intelligence capabilities are strong.
  • Figure’s videos leave him with a dual impression: “very impressive” and “a bit marketing-heavy.” After seeing a smooth demonstration of a robot tidying a living room, he wanted to move the objects himself on site to test the real performance, recalling the Tesla clip later explained as having involved a teleoperator removing a headset.
  • On product design, he likes 1X, the white, home-friendly feel of 傅利叶’s GR-3, FAUNA’s flat-headed design and the XPeng robot. These aesthetic judgments are separate from his ranking of intelligence capabilities.

24. China Has Not Missed the Intelligence Window, but Needs More Strategic Discipline

  • 程曼祺 noted that 许华哲 returned to China because he believed in “the East rising and the West declining,” while most of the intelligence leaders he now follows are not Chinese companies. 许华哲’s answer was that US founders may have greater strategic staying power.
  • He uses PI as an example: the company’s goal was embedded in its name from day 1, and it has continued pursuing the same vision, releasing progress every few months without an explicit expectation of near-term commercialization.
  • He believes more funding creates more room for error, while the industry atmosphere generates peer pressure. If a class rewards sitting up straight, everyone sits up straight; if it rewards speaking, everyone competes to speak. 许华哲 is willing to share his judgments publicly in hopes of pushing the industry toward greater focus on fundamentals.
  • He worries that Chinese embodied AI could miss “the biggest watermelon”: if China does not pursue general intelligence, the brains inside tomorrow’s metal bodies may not be controlled by China, meaning it would lose the right to define the future. But he is unequivocal that China has “definitely not missed this window yet.”

25. Once the Singularity Arrives, It May Be “Game Over”

  • 许华哲 calls one important future boundary the singularity: large models writing and improving large models, and robots assembling and manufacturing robots. Once intelligence reaches that point, the game may be over.
  • Until then, competition remains. Asked when general intelligence will fully reach human level, he offered only a rough personal estimate: perhaps about 5 years away, not a firm date.
  • Among space, quantum computing, embodied intelligence and controllable nuclear fusion—fields that could change human society—he sees embodied intelligence as relatively the most certain, even if it remains a long way off. He also acknowledges that he does not understand the other fields well.

26. Data Should Not Be Sold Easily to Competitors, and Shipments Are Not Enough

  • 许华哲 sees data as ammunition. Nvidia and other companies cannot collect it themselves, and manual operating costs are high, so they may buy data from companies with large-scale operating capabilities. Selling data to competitors to raise financing or improve the financial picture is dangerous, in his view.
  • He distinguishes the risks of domestic circulation from selling data to the US. Moving data around within China to generate revenue is understandable; selling it to the US is dangerous, and 破壳 at least “might not do it.”
  • The second problem is mindless mass production. The metric he most wants to see is robots’ daily active usage: after 5,000 units are sold, are 10, 100, 1,000 or all 5,000 actually active? Robots are physical assets, so shipments are easy to count; daily active robots are much harder to measure.
  • If 1,000 robots are sold but only 20% are used by real users, and only 20% of those are used daily, that would be disappointing. Dancing is already a mature performance; continuing to increase jump height or the number of spins has little to do with intelligence. But if athletic ability serves a specific extreme use case such as crossing mountains, it remains meaningful.

27. From Neural Networks and Game AI to Robots, the Standard Has Always Been Greater Difficulty and Impact

  • In high school, around 2009–2012, 许华哲 discussed neural networks and AI with classmates from the computer-science competition track in a computer lab, and later competed in physics. His first attempt at using a neural network to write game AI in college failed; he switched to traditional search and the A* algorithm, but his ranking was not particularly good.
  • During his PhD, he worked on a range of game AIs, including Atari, Super Mario, StarCraft and Generals.io. He also spent roughly 6 to 9 months writing StarCraft AI full time. For games such as mahjong, Texas hold’em, real-time strategy and card games, he believed reinforcement learning could eventually solve them with sustained investment.
  • He later concluded that reinforcement learning was too good at games, whose challenges and impact were limited. Robots face the physical world, where the challenge is greater and the potential service extends to everyone. By the third or fourth year of his PhD, around 2018–2019, robotics had gradually become his central path.

28. PhD-Era Doubt Rewrote the Good-Student Track into an Impact-Driven Mission

  • Once his research results, graduation requirements and advisor’s expectations were broadly satisfied, 许华哲 saw a clear linear path: publish more papers, do a postdoc, then join Google, Facebook, a Chinese company or a university. He began asking whether following that path would have any meaning.
  • His answer was blunt: “Most research has no value.” It is noise along humanity’s path of progress, with only a small share of research truly laying another brick toward the future. He wanted to do research “that isn’t noise,” and things that weren’t noise, helping more people while pursuing an extraordinary life.
  • Choosing the harder option has been a longstanding tendency. When a recommendation through the physics competition did not lead directly to the outcome he expected, he could have safely attended a university in Shanghai, or retaken the Tsinghua entrance exam and, if he failed, returned to the national college entrance exam. His parents preferred the first option, but he chose the second with almost no hesitation.
  • “Impact-driven” does not mean money or fame to him; it means improving the lives of more people. He cites microscopes, rockets, antibiotics and nuclear energy as technologies that could have enormous long-term impact, while acknowledging their potential dual use. His choice is to work to direct technology toward good.

29. A Technical Aesthetic of Simplicity, Consistency and Repetition Without Boredom

  • Music, literature and theory are all precise descriptions of the world in 许华哲’s eyes. Schubert’s D.960 gives him a sense of the divine and of heaven; Bach’s Chaconne evokes the universe unfolding and folding back in. These works also help him understand his own pain and mission.
  • Piano and tennis satisfy his preference for repetition. Practicing the same musical passage or swinging the racket again and again may look like standing still, but rhythm, angle and quality change subtly. He compares life to a spiral: from one direction it is repetition; from another, ascent.
  • Intelligence may appear to be improving at astonishing speed, he says, but at its core it is still Transformer plus data, with each local component improving continuously. Taking one problem and pushing every part to the limit suits him better than constantly changing the subject.
  • His technical aesthetic begins with “using a small method to solve a big problem”: a simple method applied to a complex problem is beautiful, while a complicated method that solves only a tiny problem is ugly. He also values conceptual consistency and dislikes combining mutually contradictory approaches. Usefulness and beauty, he believes, are almost orthogonal dimensions.

30. 2025 Brought Major Progress, but No Clear “Embodied Transformer”

  • 许华哲 believes π’s work in 2025, the scale of the Generalist dataset, various embodied VLMs and the high success rates enabled by reinforcement learning have all advanced embodied intelligence. Together, they have shown the industry the potential of data scale, semantic understanding and interaction.
  • His own reinforcement-learning project R1-00 trained separately on 7 tasks and achieved high success rates on all 7, but it did not train directly on the combined data from all 7 tasks; simply mixing them together does not immediately work.
  • He does not believe any current project can clearly be identified as an embodied equivalent of the Transformer breakthrough in 2017. Academic awards also have limitations: NeRF did not win the best paper that year but later became an important line of work. Whether a paper ultimately matters requires time and practice to determine.
  • The results he likes most include RoboCook, a system-level task spanning rolling dough, preparing filling and folding dumplings; DP3, which combines 3D vision and diffusion models; and remote tactile experiments. The latter let people in Beijing feel the softness, hardness, shape and roughness of objects touched by a robotic arm in Shanghai, and also used a breast-cancer model for remote palpation experiments.
  • Remote tactile sensing remains experimental; temperature and texture are still difficult to simulate. The team has also built the Nine Detect tactile sensor and Tactile Display, which lets people feel machine touch. 许华哲 hopes that one day people on Earth could even feel what a robot is touching on Mars.

31. Entrepreneurship Restored His Childhood Agency—and Raised the Stakes to the Maximum

  • For 许华哲, continuing to make podcasts and content is both a way to relax and a way to make information more transparent and equal. He wants people who would otherwise never encounter these stories to gain access through the platform.
  • For consumer products, social-media feedback offers rare, unfiltered market input. As the tennis world puts it, even Federer could become a 2.5-rated player on Xiaohongshu; everyone would point out what he did well and what he did badly.
  • He sees no contradiction between being outgoing and humorous and enjoying solitude and introspection. While reading 哈萨比斯’s biography, he especially resonated with the idea of a life lived to the fullest: keep moving forward like a marathon runner, and ideally cross the finish line having exhausted yourself.
  • His estimate of startup success is that each of 100 companies might have only roughly a 1% chance or less, and all 100 could also fail. What matters is whether he gave everything, created something, and whether humanity ultimately gets there.
  • He is willing to pay “all the costs,” while learning and filling in the capabilities required to build a company himself. A co-founder with a secondary-markets background will provide complementary expertise.
  • Entrepreneurship has brought him back to the 2 paths he imagined as a child: starting a company and becoming a teacher. Working from 8:30 a.m. to 12:00 a.m. every day, with a strong sense of ownership and agency, has made him feel happy again.
  • He hopes that by his next birthday or the end of the year, his child will be speaking, the company will have a strong team, they will have built something, and he will still be enjoying the process. From the “barren” office they have now, he wants to see a robot gradually “break out of its shell.”