Vol.227 The Embodied-AI Newcomer and the Queen of Investing: Innovation, Hype, Bubbles and Cycles
Vol.227 The Embodied-AI Newcomer and the Queen of Investing: Innovation, Hype, Bubbles and Cycles
Summary
- Embodied AI is still well short of the chasm. 高继洋 sees the industry still in the innovator stage, with 2-3% penetration; 星海图’s roadmap is “from developers to productivity.” The timeline is specific: robots should begin operating in real-world scenarios in the second half of this year, but “the ROI does not pencil out if you do the math on a standalone basis—that is certain”; some scenarios should work economically next year. Overseas adoption will be easier because labor costs are higher, but overseas players face a supply-chain bottleneck.
- Data is both the biggest bottleneck and the biggest uncertainty. The target is 1M hours of real-robot data this year and 10M hours next year—徐新 cited Google’s view that 10M hours should be enough to build a reliable robot capable of basic household chores. But “we’re really not sure how long it will take,” and she sees the timeline as the business’s real risk; competition is “not that concerning.” The G0.5 foundation model has already made it work: in some To B task scenarios, it “can already replace a person,” with success rates around 99%. What remains is speed, and “speed translates into economic value.”
- The business model has 3 stages, with growth accelerating at each step. First comes whole-unit sales, growing 30%-100% annually; then subscriptions for solutions in individual productivity scenarios, growing 3-10x; finally, once enough devices are deployed, selling tokens, with annual growth of dozens of times. The core logic is that “the more the brain learns, the smarter it gets,” rather than the traditional manufacturing logic that more units make products cheaper. Embodied AI also spreads differently from GPT: “If it is working in a factory, how does everyone in the world suddenly experience it?”
- 徐新的 chasm framework was the most tradeable part of the discussion. After crossing the chasm comes a 200% growth phase lasting 3-5 years, when the fastest player’s market share becomes “permanent”—provided the company becomes an operating system or infrastructure layer; the OS market cap ultimately reaches 3-5x that of infrastructure. Large language models have already crossed the chasm, at roughly 20% penetration: Nvidia is up 80%-90%, while Anthropic, Kimi and 智谱 are growing 10x a year. The main engine of future economic growth will be the token economy; outside it, even 10%-20% growth is hard-won.
- Bubbles are not the danger; the financing window is. Investors are always driven by greed and fear; greed currently dominates, with everyone gripped by FOMO. The rule is “raise money when you can”—by the time you need capital, the market may no longer be open. After a bubble bursts and others are bleeding, adding to your position is “the most economical” trade. The end state will be a 3-5-player industry, like data centers, because fixed costs are enormous—large models cost $1B a year, while embodied VLA spending is moving from RMB3B-RMB4B a year toward $1B—and the best brains are scarce.
- 星海图’s differentiation is a string of once-contrarian bets. It chose to build both the whole machine and the intelligence, hardware before intelligence, and to abandon simulation and “vegetable farm” data collection in favor of real-world data from open environments. “Every leading open-source embodied foundation model released anywhere in the world since 2026 has been trained on real data. That is a fact.” Its answer on the moat is blunt: “Iteration is the only moat.”
- AI is rewriting organizations and the founder profile. There are no middle managers, only founder mode; one founder hires based on the formula “does a person earning RMB50K a month plus an agent outperform an agent?” and believes 20-30 people could generate RMB10B in sales without a problem. Kimi has roughly 300 people and DeepSeek roughly 200. 徐新的 core advice to 高继洋 is to spend 80% of his time on the direction of the brain, reading papers and talking to the best researchers: “Your brain, your vision and your ability to mobilize resources cannot be replaced.” The underlying faith is the Bitter Lesson and the Tesla story: after abandoning 9 years of data and 300,000 lines of C++, Elon Musk closed his eyes for 20 minutes and said, “this will do”—“only a founder dares to make that decision.”
Deep dive
1. Signing a term sheet in a day: See the forest first, then pounce on the thoroughbred
- 徐新的 opening principle was direct: “When you look at a new industry, your biggest mistake is seeing the trees but not the forest—you need to examine every tree.” Around the Lunar New Year early last year, she met every company in embodied AI. She first met chief scientist 赵航 in Beijing, heard that the company had a “workaholic,” and flew to Suzhou the next day to meet 高继洋.
- There was no pitch deck at the first meeting. The office was full of standing robots, many of them “holding forth” as if they were running the company. “I said, wow, you’ve changed so many things—weren’t you a software person?” She signed the term sheet that day, then increased her position in 2 follow-on rounds. This was not standard practice, she said: “When you find a thoroughbred, pounce on it quickly.” The previous founder she had backed on the same basis was 杨植麟.
- 高继洋 acknowledged that 星海图 built its first R1 wheeled-arm body in 8 months without following the original process: “We just started with whatever was in front of us, then kept changing and building.” 徐新的 takeaway: “A young calf is not afraid of a tiger… The people who change an industry are often outsiders.” 李翔 closed the segment with a joke: apparently “Tsinghua and Peking University are no match for guts” also applies to Tsinghua graduates.
2. “Intelligence defines the body”: Build both the brain and the body, hardware first
- 徐新 remembers 2 major strategic choices from their first meeting. First, “intelligence defines the body”: the company would build both the brain and the body. Second, it would build hardware first and the brain second. “A year and a half later, looking back, it was still absolutely the right call.”
- 高继洋’s logic came from autonomous driving’s lesson: intelligence for the physical world requires physical-world data, and that requires a vehicle to carry and capture the data—the whole machine. As a software supplier, the data acquisition path had been too fragmented and “the flywheel could not turn,” which was fatal. There was no ready-made whole machine in the market. “That’s the opportunity… The only thing missing was knowledge. If we didn’t know, we could learn.”
3. What does an embodied-AI founder look like? The best brain × mass-production engineering
- 徐新 has invested in foundation models and is convinced that the premise is “the best brain.” “The brain is actually very hard to build… Where are the best brains? Tsinghua’s interdisciplinary institute.” 姚期智 brought back the best overseas PhDs to teach the hardest-to-enter Yao Class: “The best teachers teach the best students.”
- But a strong brain is not enough. “To build the body, you need engineering experience—you need to have done mass production and supply chains.” 高继洋’s team at Momenta had taken software through mass production and onto vehicles. “That combination matters because the existing pool is small while the incremental opportunity is huge. Once you have the right people, the founder combination matters enormously.”
4. How can software people build hardware? Decomposition
- 徐新’s pointed question was: “Why can you build hardware?” Hardware is traditionally the domain of veteran practitioners, so how would a group of young people solve it? 高继洋’s answer was an engineering mindset: “For any complex problem, we decompose it. Break it into a number of sub-problems, examine each one separately, find the right person for each, solve the concrete problems one by one, then aggregate them into the overall solution.”
5. Why partner up? Too big, too hard, too disruptive
- 徐新 asked 赵航: each of you could raise a great deal of money on your own, so why join forces? His answer stayed with her: “This is an especially big and difficult undertaking. One person doing it alone has a lower probability of success; several people doing it together raises the odds materially.”
- The 3 founders share a common starting point. 高继洋 at USC and 赵航 at MIT worked on computer vision during their PhDs and read each other’s papers. Both joined Waymo in 2019: “We all wanted to build robots, and autonomous driving was the industry closest to robots at that time.” 天威 was 高继洋’s colleague at Momenta after he returned to China, also coming from intelligent driving.
- 高继洋 is candid about his own role: “赵航 has a better tech vision than I do… I’m more of a generalist. I can talk to the best scientists, but I don’t have the best scientists’ tech vision. A CEO who comes purely from industry can’t talk to scientists—their thought processes and logic simply don’t line up.”
6. The chain is longer than expected: From gear backlash to current throughput
- “The chain is longer than we initially thought.” What looked like 2 pillars—whole machine and intelligence—turned out to include data behind intelligence, and scenarios, operations and equipment behind data, with storage, transmission and other infrastructure still ahead. “Large language models have a huge parameter count, but not that much data—it’s all text. We work with multimodal data, and the data volume is enormous.”
- Looking up the hardware stack, the entire supply chain remains immature. Power modules break down into frameless motors, reducers, and gear machining, where backlash and consistency are major issues. Batteries must handle the current draw when every motor in the body surges simultaneously; “it’s not like a phone with stable output.” Then come edge compute, material density and total system weight. “There are so many things to optimize. That’s why courage matters: don’t change the big strategy, and improve as you learn.”
7. Scientist-founders are scarce: Tech vision determines the outcome
- 徐新’s observation is that among 10 scientists who say they want to start a company, very few actually do. “After finishing a PhD, they are rational and the opportunity cost is high. Many are pursuing truth—wherever AGI is, that’s where I want to go.” Even a star like Karpathy chose to be someone else’s number 2 rather than run his own company. “Of 10, maybe 0.5 want to start a company, and are both excellent and creative. That is a scarce resource.”
- Students follow the technical leader, not the equity package. The definition of technical excellence is tech vision: “seeing what others cannot see.” Her example is OpenAI: Ilya was “the true inventor of large language models,” and after he left, “they lost something on tech vision.” The bet has to be right. After burning tens of millions of dollars or several hundred million RMB, “you can’t have everything at once.” Anthropic bet on a single direction—and got it right.
8. Resolving disagreements: Hardware first, at 赵航’s expense
- 高继洋 does not dodge conflict: “Disagreements objectively exist. The key is how you resolve them.” The prerequisite is aligned goals and values. The hardest example was the 2024 decision to build the whole machine first. “In hindsight, the decision was right. But if you go back to that point in time, it meant sacrificing 赵航—hardware had nothing to do with him, and there was less for him to do at that moment.”
- The reverse applied when the company shifted toward intelligence. 高继洋 told the whole-machine team not to treat external customers as the top priority; it had to prioritize the internal customer, 赵航, even more highly. Everyone subordinated themselves to the more important mission. 徐新的 summary was simple: “Find the right people and everything falls into place. Get the people wrong and you spend every day resolving conflicts.”
9. A 5-year robotics investing lesson: Spring arrived last year
- Capital Today began investing in rule-based robotics 5 years ago, including 高仙 cleaning robots, 海柔 warehouse systems, XYZ loading and unloading, and Vision Nav autonomous forklifts. That was “a roughly 2-year course of study.” 徐新 believes spring arrived last year: hardware became robust, most corner cases had been worked through, and aging populations plus rising wages changed the economics. “It can’t compete with people on efficiency, but it can do the dirty, difficult and exhausting work people don’t want to do.” Installing solar panels in the desert is one scenario that has worked.
- She also acknowledges that hardware is hard, recalling a conversation with Peter Thiel. Elon Musk was far more capable than Mark Zuckerberg, but after more than a decade Tesla had less than 10% market share; Zuckerberg “quickly monopolized” his market. “Bits scale easily; atoms are hard. There are only a few supply chains, and you don’t have the final say.”
10. The chasm is still ahead: The current 2-3% is the innovator phase
- Both men had read Crossing the Chasm. 高继洋 mapped the stages: innovators at roughly 2-3%, early adopters at about 13%, and a chasm around 15%-16%. “If you can’t cross it, you fall to your death.” The industry “definitely hasn’t crossed it yet; it is still in the innovator phase.” That is why 星海图’s business logic is “from developers to productivity”—the journey from innovators to the majority.
- What is missing is the shift from a technology model that is merely usable to one that is useful. His sense of progress is that foundation models improved over the past year “faster than an infant grows.” Robots can now rummage through bags, pick up parts and grasp all kinds of objects at a basic level. “The problem now is that they aren’t fast enough. The success rate today may be 99%.”
- He also praised the training grounds and pilot-scale testing facilities being built by the government. “Building a bridge between final use and R&D is really building the bridge across the chasm. That is exactly the right thing to do.”
- His timeline was unusually direct: robots should begin operating in real scenarios in the second half of this year, but “the ROI does not pencil out if you do the math on a standalone basis—that is certain. In some scenarios, it will work economically next year.” Overseas adoption will be easier because labor costs are higher, but the bottleneck for overseas companies is not price; it is that “the supply chain cannot keep up with their pace of progress.”
11. The data lesson: From 1M hours to 10M hours
- 李翔 laid out the gap clearly. LLM pre-training data is already available across the internet; autonomous-driving data was collected automatically as the cars drove, with users doing the collection for free. Robot data has to be collected from scratch: teleoperation, grippers and all the rest. “We have to make up that course.” The devices are also entirely new and expensive. “Both of these have to be rebuilt from zero.”
- 李翔 repeatedly pressed for a quantitative anchor. The team is confident about 1M hours this year and reaching 10M hours next year. “Google’s view seems to be that 10M hours is basically enough to produce a reliable robot, one that can handle some household chores.” But the time required to get there is uncertain, and “the entire methodology and technical process may have to change.” 徐新的 optimism is rooted in China’s hardware supply chain: “Once the iteration is complete, maybe this whole thing can be done in 2-3 years.”
12. Embodied AI will not spread through a GPT moment: Quiet penetration through To B
- 高继洋’s view is worth recording: “Embodied AI does not penetrate the market the way GPT did. GPT produced an extraordinary product overnight; people downloaded it onto their phones, social media spread it, and most of the world experienced it. But when embodied AI is being used, it may be working in a particular factory. How does everyone in the world suddenly experience it?”
- The first stage is productivity and therefore To B: “solving staffing problems in selected task scenarios, rather than truly replacing a person.” For To C, he sees companionship, entertainment and emotional value. “That is not productivity.”
- 徐新 added context on the agent adoption curve. After the Open Crawl agent appeared in January this year, “token consumption was 100x ChatGPT’s, and eventually it will be 1,000x.” But that happened because the endpoint and distribution channels were ready. Embodied AI is still on a linear curve because “the device is not ready.”
13. The 3-stage business model: Growth accelerates at every step
- 高继洋’s full outline: Stage 1 is R&D-and-manufacturing-oriented whole-unit sales, growing 30%-100% annually. Stage 2 is building solution capability in a single productivity scenario and selling subscriptions, with 3-10x growth. Stage 3 is selling tokens once multiple scenarios work and enough devices are deployed, with growth of dozens of times in a year. “The embodied-AI industry could very plausibly grow faster and faster.”
- What 徐新 most agreed with was the scale-effect framing from that day’s presentation. The brain learns and becomes more valuable the more it is used; hardware becomes cheaper and less difficult to build at scale. That is why both sides must be built. She cited the broader pattern from the book: technology penetration takes 10-15 years, product life cycles last 20-30 years, and Microsoft has continued to collect the dividend since the PC era.
14. 徐新的 chasm economics: Market share in the hypergrowth phase lasts forever
- This was the most systematic section of the discussion. After crossing the chasm, growth reaches 200% for 3-5 years as the 34% early majority enters, costs fall—“special thanks to DeepSeek for cutting costs 10x”—and applications proliferate. There are usually 3-5 players. “The person who moves fastest and gets big fast takes market share that lasts forever.”
- That permanence has 2 prerequisites: the company must be an operating system or infrastructure. Windows and Intel in the PC era, and Qualcomm plus iOS/Android in mobile, were all monopoly-like businesses. She once debated whether large models were operating systems. “I don’t debate it anymore: the OS and the agent are one. A huge ecosystem can form a monopoly.”
- The sequence is equally tradeable. Infrastructure rises first—Nvidia’s AI chips are up 80%-90%—but operations then move faster. “How fast is Anthropic growing? Kimi and 智谱 are growing 10x a year.” At the end state, the OS market cap is 3-5x infrastructure’s. She believes large language models have crossed the chasm, with roughly 20% penetration; 豆包’s MAU has reached something close to the entire internet’s scale.
15. The token economy: Outside it, growth becomes painful
- 徐新的 macro framework: “One of the main drivers of future economic growth will be the token economy—creating wealth by producing tokens. If you are not on it, growth becomes weak and defensive; even 10%-20% gains are hard-won. If you are in it, 200%-300% annual growth is not in doubt. As a VC, it becomes obvious where to put the money.” Asked what happens to retail consumption, she said it is not over: supply chains and private labels can still produce “maybe 10%-20%” growth.
- That led to her Google-versus-Anthropic debate. Most people would invest in Google: 黄仁勋’s 6-layer cake has 5 layers, so the risk is low. “But growth is only 20%.” Anthropic has no legacy burden and only one large model. The most attractive asset in the value chain is clearly the large model itself: the company doing only that is lighter, more focused and more aggressive.
16. 3 sentences unchanged in 3 years; the path has changed countless times
- 高继洋’s 3 founding statements have not changed: “The future of embodied AI is one brain, many tasks; the core is the brain, not the form. The long-term moat is a closed loop of physical-world data. The key path to building that loop is the whole machine plus intelligence.” What they had not figured out was execution: which product to start with and which market to enter changed repeatedly. 徐新 confirmed that this was the message in their first meeting and “has basically not changed.”
- Each stage also brought a then-contrarian choice. Building the whole machine before the intelligence in 2024-25 drew 2 criticisms: “Get the demo out quickly; without a demo, your intelligence doesn’t work,” and “You can’t beat specialists at building the whole machine; you should focus.” “We stuck to our judgment.” Data collection followed the same pattern. In 2025, many companies collected data in “vegetable farms”; 星海图 did not. It collected in “open environments” to prove the data was higher-quality and effective. Good hardware plus good data meant “my foundation model was good.”
- Is real-world data now a consensus? He refused to answer in terms of consensus. “Consensus is subjective. We’ll use facts: every leading open-source embodied foundation model released anywhere in the world since 2026 has been trained on real data. That is a fact.”
17. The lesson from missing ByteDance: Study winner patterns early
- 徐新的 3 learning methods are: learn from founders—she does a deep dive with 高继洋 every 2-3 months, asking only 3 questions: What have you learned recently? What has been painful in your growth? What is the moat at the level of first principles?—study winner patterns, and develop consumer insight.
- Her biggest regret is ByteDance. 张一鸣 asked for a high price—“$7B, or whatever year it was,” she no longer remembers—and she was scared off. “Later I realized I hadn’t understood it. You really cannot make money beyond your level of understanding.” She understood it after reading all of Mark Zuckerberg’s quarterly earnings-call transcripts. In 2012-13, he said all information would eventually be expressed through a feed and ads would be embedded in the feed as seamlessly as content. “That was the moment I knew how powerful a recommendation engine could be.” She now goes through every speech and podcast by the founders of the most advanced labs.
18. 7-Eleven’s sushi: Human laziness and the heavy bet on Meituan
- In the mobile-internet era, 55% of one fund went into Meituan. The key question was whether users would keep ordering after subsidies ended. 王慧文’s real-world test was inconclusive: “Once the subsidy made it cheap and we took it back, orders immediately fell by 20, and we lost our nerve.”
- The answer came from a biography of 7-Eleven’s 鈴木敏文. After the war, he sold sushi made with the best rice and deep-sea fish at half price. Asked whether people would still eat it once the discount ended, he said: “Human nature is lazy. Once you get used to eating something, when someone else has prepared it for you and you gradually raise the price, they will keep eating it.” “The decision to make a heavy bet on Meituan came from that judgment about 鈴木敏文: don’t offer the delivery subsidy once; offer it 5 times, build the habit, and users won’t go back.”
19. Judging people: Killer instinct and the ability to see what others cannot
- Since she began investing in 1995, 徐新 says the first common trait among founders is killer instinct. 杨植麟 said very early that he wanted to build long context and was among the first to work on agents. When people at Alibaba said they could not find a competitor to 刘强东 “even with a magnifying glass,” he went the other way—building a heavy logistics network and a delivery operation in every city. At 20 orders a day each city lost money; break-even came at 2,000 orders and would take a year. 80% of customer complaints came from delivery.
- Her message to young founders follows directly: “AI can rebuild every industry, just as e-commerce washed through every retail category. Founders who understand AI will disrupt founders who do not. That will be the growth engine of the next 10 or 20 years.”
20. There is only 1 moat: Iteration
- 李翔 said it was harsh to ask a 3-year-old company about its moat. 高继洋’s answer was unequivocal: “In an industry where innovation happens every day, iteration is the only moat. Think you have something? Stop for 2 days and you’re behind. People ask what 星海图’s DNA is—is it the DNA of building brains? Our DNA is iteration.”
- 徐新的 early-stage version is consistent: high ambition, patience for the long term, good killer instinct, fast learning and exceptional execution. “At the early stage, that is the moat. This is a huge undertaking with no product-market fit yet. You are looking for reliable founders in a sufficiently large and disruptive arena.”
21. The AI-era organization: No middle managers
- 徐新 meets 4 or 5 founders a day and sees a new species: all of them are in founder mode, with one person directing 3-5 AIs like a general taking a city, sleeping only 4-5 hours a night. “It isn’t that they are overworking; agents keep pushing the human limit upward.” Her new diligence metric is to first see whether the boss is exceptional and whether the company has burned 100M tokens.
- The organizational formula has changed. One founder hires only if “a person earning RMB50K a month plus an agent is greater than an agent.” Otherwise, the candidate is rejected; the company seeks people earning RMB100K-200K a month who can direct agents. “Organizations really will have no middle managers in the future. Maybe 20-30 people can generate RMB10B in sales without a problem.” Kimi had only a little over 300 people at the end of last year; DeepSeek had around 200.
- 高继洋 uses GPT for deep discussions about patterns. “I get facts from the front lines. But patterns have happened before—who can compress those patterns extremely effectively and put them in front of you? GPT can do that.”
22. The ritual of deep thinking: Planes, meditation and the major question
- 高继洋 sleeps 7-8 hours and exercises consistently. “Don’t fall into strategic laziness and tactical diligence.” His setting for deep thought is an airplane: no interruptions and no internet. He takes one problem he has been thinking about, opens a notebook and writes keywords. “Many times, the moment I get off the plane I start calling colleagues. The insight arrives, and I redeploy.”
- 徐新的 version is “do fixed things at fixed times.” Every morning after meditation, she reviews what she learned the previous day. On weekends she thinks through major questions: “What is my major question this month? Maybe I can’t answer it for 1, 2 or 3 months, but it stays on my list. Major questions require repeated, deep thinking.”
23. 3 kinds of intelligence: Instinct, task and evolution—a name GPT supplied
- The roadmap was born from a conversation in a café: beyond VLA and WAM, what else is there? They first defined instinct intelligence: “Naturally commanding your own body—2 arms and 2 legs today, but could you command 4 arms? It is somewhat like the cerebellum.” The term came from 赵航’s student 庄子文’s Project Instinct. Pushing further, they asked why the human body evolved this way. The answer was evolution: “So why can’t AI do the same?” That produced the 3-part framework of instinct, task and evolutionary intelligence.
- The footnote is that they had not initially thought of the term “evolutionary intelligence.” They described the rough idea to GPT, and GPT suggested the phrase. 徐新 only heard the outline recently, in a conversation over the Lunar New Year. “This is the first time it has been explained in this much detail.”
24. The 5-piece puzzle: Talent, data, capital, supply chain and distribution
- 高继洋 breaks the conditions for realizing the vision into 5 pieces. Talent is “the starting point of everything”; this year the company plans to build a joint embodied-intelligence research center with Tsinghua University to bind high-quality students earlier. Then come data, through the Yizhuang data program; capital, which is “being solved reasonably well”; the whole-machine component supply chain; and distribution. “Embodied AI does not yet have a distribution channel. Large language models were distributed through social media and phones. We have to build ours from scratch.”
- The supply chain contains a structural opportunity. Over the past 3 years, a large number of automotive-grade component companies have entered the sector looking for a second growth curve. “Their capabilities, qualifications and scale are better and larger. The momentum and potential of this new chain are higher.” He describes suppliers’ mindset as “aspiration plus anxiety”: they are willing to iterate with a startup “because everyone is looking at the sector, not us. We just happen to look like the leader.”
- “Squeezing all 5 dimensions into one startup is complicated. But that is exactly what is good about it: it is so complex, with so many elements and such a long chain, that large companies do not have a particular advantage. This is a sector naturally suited to entrepreneurs.”
25. Why large companies cannot break in: Outside their range
- 高继洋’s full argument is that the most dangerous threat to a startup is an incumbent’s existing business synergies and dimensional attack. But embodied AI “is not within any large company’s direct range.” On the supply side, foundation models did not previously exist, and neither did the data; everyone starts from the same point. The closest whole-machine capability is automotive, but “which components can be used directly? None. And if you bring automotive manufacturing logic to it, the problem is that you move too slowly.” On the demand side, there is no existing terminal or traffic monopoly, and no one monopolizes labor. JD.com, the largest private employer, has roughly 1M employees—a small share of the labor market.
- He has also run the capital comparison. Xiaomi and Huawei have several hundred billion RMB in cash reserves; mid-sized companies have tens of billions. “But they have a core business to run. How much can they actually focus on this one thing? Maybe less than we can. We have more focused capital, higher organizational efficiency, and no business synergies to distract us. We should win.” Leading embodied-AI companies already have tens of billions, in some cases high tens of billions, of RMB in hand. “This is an era that rewards innovation.”
26. The only competitor metric that matters: 12-month rate of change
- 高继洋’s competitive dashboard has only 1 metric: “I look at how much it has changed over the past 12 months. Even if it is behind us today, if its rate of change exceeds ours, I have to pay very close attention—it means it is iterating faster than we are. Large companies iterate slowly; if they are not committed to investing, there is not much to worry about.”
- 徐新 used Google to show both the danger and the slowness of large companies. Google created the Transformer, but not ChatGPT. It took almost 3 years to fix the organization: bring the founder back, make Hassabis the sole leader overseeing DeepMind and Google Brain, and provide enough TPU capacity. “It immediately won its first battle and got the banana; Gemini went into overdrive.” The Chinese comparison is ByteDance: it can adjust its organization in 1-1.5 years rather than 3, has huge cash reserves, and can make miracles through sheer force. “In the long run, it will definitely be a formidable competitor. But it is not focusing on these areas now. That is our time window.”
- The counterintuitive consensus is that slower foundation-model progress is not necessarily bad. “If progress opens everything up too quickly, all the large companies will come down.” The key is where you are when the inflection point arrives. Get out in front, and large-company entry can actually make the game more interesting: more resources come in, and you can access more of them.
27. Bubbles are nothing special: Raise as much as possible while greed is in control
- 徐新 has lived through multiple cycles. She invested in NetEase, Sina and Sohu in 1999; they listed at the bubble’s peak in 2000, then the bubble burst in the second half of the year. “Blood ran in the streets. Every internet company died.” But those 3 companies had raised $80M-$100M during the bubble and survived. They had “3 years without competitors,” enough time to build a business model. Under pressure, 丁磊 turned to games and later became China’s richest person.
- Her operating discipline follows from that experience: “Investors always have 2 states of mind: greed and fear. When they are fearful, nothing you say matters; they don’t believe you. When they are greedy, they will invest at any price because of FOMO. Everyone is greedy now. So list if you can and raise if you can. Raise money when you can, not when you need it. By the time you need it, the market may have changed and you may not be able to raise.”
- On bubbles themselves, she echoed the Bill Gates-style view: “In the short term, people overestimate it; in the long term, they underestimate it. This is a sufficiently large opportunity, and it is not actually that expensive.” When the bubble bursts and washes out competitors, “you can keep adding to your position. That is the most economical trade.”
28. A post-bubble playbook: Absorb the 200 people and pick up rule-based assets
- 高继洋 has rehearsed what happens after a bust. “The direction of an industry in China may ultimately be determined by 100-200 people. If they consolidate into 10 groups, I have 9 competitors. If they consolidate into 5 groups, everyone else is scattered. When the bubble bursts, we should absorb these people.”
- The second move is M&A. Traditional rule-based robotics companies will be in worse shape—not because they have no value, but because capital will assign them lower valuations. Their customer channels, delivery capabilities and accumulated engineering expertise are valuable assets. He added a note of restraint: “We will read the situation and act accordingly, not attack recklessly.”
29. AI logic versus manufacturing logic: Every run is non-standard production
- 高继洋 manages the financing and spending paths separately. On financing, “take the money when it is available.” On spending, the first driver is the scaling law. “The AI mindset is: I’m at 1 now, and I have to see what happens at 5 or 10. Even if I don’t know what will happen, history proves that 5 and 10 always bring surprises. The market rewards those surprises.”
- Traditional robotics uses a different model: this robot costs RMB50M to develop, so how many units—1,000 or 2,000—are needed to earn it back? “Apply that logic to AI and you’re finished; you won’t dare invest because you can’t calculate it.” The biggest obstacle for traditional robotics companies entering embodied AI is whether they can invest in the AI way. The cycles are also completely different: traditional robotics has 18 months of R&D, 3-4 years of sales and 2 years of after-sales service. AI needs a 3-year base R&D timeline, and a large model’s “every production run is non-standard production,” while a robot becomes standardized and mass-produced once its configuration is set. The ROI calculation period is fundamentally different.
30. Is a large model a good business? Like data centers, only 3-5 will remain
- 徐新 applies Buffett’s test: a good business achieves relative monopoly and pricing power. Her breakdown is that large models have very clear economies of scale. Pre-training shows no obvious network effect, but post-training and RL show some. The key is coding: “The smartest engineers train it on the hardest problems, it produces results with very high intelligence, and more people use it. That is a network effect.”
- The industry will converge because of 2 forms of scarcity. Fixed-asset investment is enormous: “Without $1B a year, you are not even at the table.” Talent is scarce as well. “It is somewhat like data centers: 3-5 companies. To have only 3-5 competitors serving a world this large is actually a very good business.” The US is basically down to 3; Meta has not produced a model of its own.
- She also outlined the path to profitability. Pre-training is fixed cost; inference is variable cost and is “getting closer and closer to making money.” Everyone is raising prices, chip costs per token are falling, and model gross margins are not bad—“around 40%.” Some companies were already profitable in Q2. Open source is not necessarily bad: if agents drive token consumption to 100x chat and eventually 1,000x, “we can make it up in volume. Sending intelligence to the entire world at thin margins is a good business.” Embodied AI has the same structure: huge fixed costs are a moat, and adding the best brains makes it the hardest business—and the best one.
31. VC economics: Home-run intensity matters more than frequency
- You cannot sit out a giant wave. “Whether to invest is a question of kind; when to invest is a tactical question. When a huge wave hits, you have to jump into the water. If you didn’t invest in the internet, mobile internet or large models, you are not a very good VC.” She quoted 王兴’s analogy: traditional investing is mountaineering—the mountain remains; technology investing is surfing. Miss the wave and you wait for the next one. The method is to “jump in first with a smaller check, then add, add and add when you see signs of life.”
- A home run is a company that returns the entire fund or generates an 8x return. It has 2 dimensions, frequency and intensity: “Intensity matters more than frequency.” Frequency is luck; a fund’s 4-year investment period typically produces 3-5 great companies, with 4 as the average. Investing in more companies and hiring more people can make results worse by diluting time and attention. Intensity is controllable: double down and triple down on winners, hold long enough—“our long enough is 8-10 years”—and earn the compound return on time.
- She also exercises discipline on fund size. Invest $120M in each of 4 home runs, and $10M in each of the 16 non-home runs among 20 investments; the fund comes to roughly $600M-$800M. A 5x-6x return would be good performance. “Bigger is not useful. There are not that many great companies.”
32. The autonomous-driving lesson: The Bitter Lesson is a faith
- 高继洋 is unequivocal in comparing the 2 waves. Autonomous driving is much smaller in industry scale and social impact—ultimately “a small industry with limited influence” that cannot affect every sector. Embodied AI’s upstream is also its downstream; it will become a central industry. The more important distinction is DNA. The leading autonomous-driving companies in the first wave did not start as AI companies; they were robotics companies focused on rule-based systems and solving corner cases. The embodied companies that positioned themselves correctly “see themselves as AI companies,” solving problems through scale, foundation models and data.
- He describes himself as one of the first PhDs trained in deep learning—he entered his program in 2015, while 赵航 entered in 2013—and says The Bitter Lesson shaped him: “Believe in scale, believe in data, believe in compute. Do not believe in rules designed around human bias. Those work in the short term and fail in the long term.” Autonomous driving’s true legacy for him is a handful of foundational decisions: whole machine plus intelligence, the data closed loop and reliance on real-world data. “To be frank, there is almost nothing in the specific algorithms and technology that can be carried over.”
33. Elon Musk closes his eyes for 20 minutes: “this will do”
- 徐新的 Tesla story was the most dramatic footnote of the episode. Tesla’s autonomous-driving system was once rule-based: “When it encountered a question, the team wrote another piece of code. The C++ grew to 300,000 lines and the team was desperate.” A young engineer used a self-learning approach to build a demo on the road outside Musk’s home. After Musk tried it, the entire group debated whether to abandon 9 years of data, 300,000 lines of code and several billion dollars. “He closed his eyes and thought for 20 minutes. The room was silent. When he opened them, he said, ‘this will do,’ and switched.”
- Her conclusion goes to the nature of organizations: “Only a founder dares to make that decision. It is hard to attack your own work, and ordinary people cannot bear to discard what they have accumulated.” That is also why she insists on investing in founders born in the 1990s. “This is a dividing line: do you believe in scaling or not?” One company’s CFO even had a slide calling for the hiring of “post-95 workaholics.”
34. This time really is different: 10x in a year, replacing people outright
- 李翔 cited Marc Andreessen: the 5 most dangerous words in investing are “this time is different”—but this time really is different. 徐新 agrees, using speed as the benchmark. JD.com grew 3x in a year; food delivery grew 3-5x. This time it is “10x in a year. No product has penetrated faster than ChatGPT. Speed determines the ceiling. It is a completely different animal.”
- The fundamental difference is even starker: the internet and mobile internet were tools and connections. This time, “it is directly replacing people.” Her own experience is evidence: “The more I use GPT, the more I feel it is smarter than I am. It even seems to have human compassion. Give it a few more years and we may really be inferior to it.” She cautions against extrapolating that speed to embodied AI: task intelligence is still a long way off, with 20,000 components and a supply chain that makes the whole problem increasingly complex.
35. The JD.com retrospective: Get big fast; losses are not the problem
- 徐新 walks young founders through her best-known case. When she invested in JD.com, it had only RMB50M in sales. 刘强东 asked for $2M; she gave him $10M. The money funded 2 decisions: expand categories—Dangdang had 2M users versus JD.com’s 20,000 but sold only books, while 刘强东 believed the key to retail was abundant selection—and build delivery in-house. Each city would have 20 people and 2 vehicles, losing money at 20 orders a day and reaching break-even at 2,000 orders after a year. “We weren’t going to 1 city. We were going to 30.”
- Growth was written into the contract. An 18-point ESOP was tied to performance. 刘强东 committed to 200% annual growth; 徐新, worried he would miss, cut the target to 100%. “He actually did 200%: over 4 years, RMB50M, RMB300M, RMB1B, RMB3B, RMB10B.” The company was losing money throughout.
- Her formula still holds: “What makes you the most attractive player? You take the largest share of the market in incremental growth. Then you are the most attractive player.” Losses do not matter as long as cash flow is positive. The cash came from supplier payment terms, following the Gome and Suning model. “Before JD.com went public, it had $1B on the balance sheet. It was like buying insurance with someone else’s money.” But the rule has a boundary: “It only works while cash flow keeps coming. When greed is in control, investors will accept several more years of losses. When fear arrives, they may not let you lose for that long.”
36. The bottleneck is the timeline, not competition
- Asked what constrained 星海图’s roadmap, 徐新的 answer was not the large companies. “It is mainly the timeline: how long will it take to go from 1M to 10M hours? The methodology and process may have to change. I’m not that worried about competition. Large companies may not have an advantage. Can you actually build a foundation model?”
- The real risk is cadence. “Move too aggressively and take too big a step, and spring arrives but you have already fallen. Move too slowly and you are no longer at the front; investors stop funding you, talent concludes that the company is silent, and a vicious cycle begins.” She cited Jack Ma: “The best team, money you cannot spend and the biggest market will always produce a winner.”
37. Best and worst cases: 3 dividends and the risk of wealth monopoly
- They each laid out their tails. 徐新的 worst case is “total failure in this wave—not backing a single good company.” The best case is backing a company that changes the world. She believes China has a bigger opportunity in embodied AI than the US: China is catching up in large language models, but may lead in embodied AI. The edge comes from “the founder dividend, the engineer dividend and the government-execution dividend”—strong supply chains, strong data-organization capabilities and a government that “understands the sector extremely well,” as it did with solar.
- 高继洋 says the company will not die. “My decision-making is becoming more conservative. My risk is not dying; it is missing a major opportunity because my experience does not match the moment.” He cites OpenAI’s failure to find coding and Anthropic’s resulting lead. At the industry level, he offers a rare double-tail view. The best case is “everyone genuinely gains freedom and works for their interests.” The worst is that AI monopolizes physical-world productivity. “I’m not sure what terrifying consequences that kind of wealth monopoly would bring.”
- 徐新 takes up the social question. AI “seems to be widening the gap between rich and poor—the strong get stronger.” A form of AGI-driven equality “will definitely not arrive quickly, maybe in 10 years,” and the intervening period may bring the “pain of unemployment.” Society will need methods to help people start second careers and guidance on philosophy and values, “but no one is thinking about it yet.” Her consolation is AIGC: music, video, novels and short dramas create many things to do. “They do token manufacturing; we do consumption.”
38. Spend 80% of your time on the brain: Alchemy requires a master’s touch
- 徐新的 advice to 高继洋 was unusually direct. “A large model is not science exactly; it is somewhat like alchemy, where a master’s touch matters. You are a PhD. You should read more papers and talk every day with the most insightful people. Build systems quickly and delegate supply chain, HR and expense reimbursement; if delegation does not work, replace the person. All of that is replaceable. But your brain, your vision and your ability to mobilize resources cannot be replaced. 张一鸣 spent 50% of his time on this even when Douyin was already a huge business. For a company as small as ours, you should spend 80%.”
- 高继洋’s learning theory follows the same line: “Our lack of understanding comes from acquiring wrong information through wrong channels.” He emphasizes first-hand information from the field, followed by the loop he repeats often: information produces understanding; understanding produces strategy; strategy drives execution; execution produces results; results generate new information and feedback. 李翔 summed it up: “You’re training a model.”
- 徐新的 own information sources include long-form podcasts whose name sounds like “next freedom” and All-In, often 3-5 hours long. She first pulls the transcript and creates a key summary before deciding what to read in depth. She has carried her habit of conducting roughly 2,000 user interviews a year into To B: asking Kimi users when they use a model whose name sounds like “Anthropic” and when they use Kimi—“the company pays for this model; I moonlight with Kimi”—and using deep interviews to find the pricing sweet spot for portfolio companies: RMB90, RMB190 or RMB590.
39. The end state, profitability and founder IP: A few companies survive; those who align words and actions remain
- The 2 speakers are aligned on the end state. 高继洋: building the brain is not easy, economies of scale are clear, and “concentration will be high.” The market may end up with 4-5 large-model startups plus 2-3 incumbents; embodied AI will be similar. Capital consumption cannot support many more companies. 徐新 puts the numbers on it: training a VLA model costs RMB3B-RMB4B a year at the outset—$500M in her framing—and ultimately “$1B a year” will be required. 高继洋 declines to give a profitability timeline: “If there is room for more than 10x growth in every 3-year period ahead, we should put all available capital to work now for maximum growth. The question today is cash flow. As long as it is positive, the company is financially sound.”
- Founder IP matters in the AI era. 徐新 believes leadership cannot be invisible, but founders do not need to chase traffic every day. “The people you are trying to attract are knowledge workers, and they will find you quickly. One conference a year to share the ideal, the conviction and the product is enough.” In the AI era, “it seems no one advertises. If the product is exceptional, it speaks for itself.” 高继洋 adds an industry concern: “There is a lot of noise in this industry. Good companies without enough voice bandwidth get drowned out, and then good talent and good partners cannot find you.”
- How do you distinguish a gold miner from someone with a mission? 徐新的 answer has 2 parts: “Words and actions must align: what he does must match what he says. Then look at whether the co-founders he attracts are a like-minded group. You do not attract an excellent co-founder through personal charisma; only a grand dream can do that.” 高继洋 has no sense of superiority toward the previous generation: “I respect everyone who has delivered results. The difference really does not come from anything else; it comes from the times making the hero. If I had been born in their era, I could not have built what we are building today. Starting Meituan or JD.com was harder in some ways: anyone could do it, and competition was more intense. Our sector filters out a large number of people at the start.” The questions he still wants to ask 刘强东 and 王兴 are: what was their state of mind as the industry scaled rapidly, and what mistakes did they make? “That information does not circulate in this world.”
- The final value on the wall is “thrift and frugality,” which 高继洋 translates into lean operations: “If RMB1 solves the problem, there is no need to spend RMB2. Without detailed accounting, operating costs can be 50%-80% higher. If $100M can train 10 times, you can only train 5.” 徐新 points to the Meituan founder, who traveled on business staying at Home Inn and flew economy class. “That is a good quality.” As for pragmatism versus bold innovation: “Look at which organization you are applying it to. When an organization is full of innovative energy, make it a little more pragmatic. When it is lifeless, innovate boldly.”