戴雨森 on Harness, the Next ByteDance, and 2026's Big Opportunities
戴雨森 on Harness, the Next ByteDance, and 2026's Big Opportunities
Summary
- 戴雨森 revisits how his “year of R” call got humbled. He was right about OpenAI’s consumer ceiling: the $20 subscription is hard to raise, 50M paying users have already filtered out everyone willing to pay for a chatbot, and “everyone was too optimistic about putting ads in ChatGPT”; he was wrong about the threshold discontinuity in coding from Claude 4.5 to 4.6. His explanatory framework is worth remembering: “You only get a steam engine once the water boils; at 99 degrees, you still don’t” — intelligence creates value through discontinuities, not gradual improvement, and even Anthropic calling the model 4.5 rather than 5 suggests the company itself did not anticipate the jump.
- The return problem has not been solved; it has merely been pushed down to Anthropic’s customers. Tokens are customers’ inputs, not their output; the chain is input → output → result, and “the stagnation of mobile internet was not caused by a shortage of programmers, but by people not knowing what to build.” The numbers are now too big to ignore: AI hardware profits will reach $700B this year, Samsung and SK Hynix are approaching Nvidia’s profit level, and Anthropic’s year-end AR expectation of $100B implies selling $10B worth of tokens every month. The valuation framework has 3 parts: in the short term—3 months—both OpenAI and Anthropic are undervalued, and a public listing “could be hyped to $3T”; over a 1-2 year horizon, both are overvalued; over 10 years, the winners could be $5T or $10T companies.
- Harnesses are the OS; models are the CPU. Users are loyal to the harness, not the model: people keep rescuing the memory in OpenClaw while casually swapping in Kimi K2.5, “90-point performance at a 20-point price”; Cursor training Composer on data generated by its wrapper proves that “without the wrapper, there would be no such model.” “The model is the product” and “the wrapper is worth more” can both be true.
- His most exciting new thesis is network effects between agents. Once a harness has accumulated personalized context, different people’s agents will produce different results on the same task: “My agent hires 张小珺’s agent. $1,000 is the value of the tokens; $9,000 is the proprietary knowledge accumulated by your agent.” That proprietary information may never be distilled into the model, and agent-to-agent marketplace companies are already emerging.
- The next ByteDance will not look like ByteDance. AI feeds are about “wrapping new technology in the shell ByteDance is best at, but beating yourself under your own game rules is extremely difficult.” Consumer entertainment has to be more fun than Douyin and RedNote from day one, which is a hard problem; productivity is the “easy problem.” The opportunities are hiding in attributes that incumbents dismiss: wrappers, open source, and insecurity. “If you think wrapper value is zero, what if you’re wrong?”
- His positioning and investing approach: after exiting his positions before the Lunar New Year, he added back “chokepoint” hardware exposure in storage, optical components, CPUs and related areas, but remains unwilling to be aggressive. His investing idol is Stanley Druckenmiller; he trades like a trader, with a “strong opinion weekly held” mindset. In venture, he continues to back people rather than chase themes: 刘松明 and 丁宁’s embodied-AI companies have risen from valuations of roughly RMB200M to several billion renminbi. His consumer-electronics hot take: AI hardware will repeat the fate of new consumer brands, and wearable recording devices are a “need invented by VCs.”
- At the individual and social level, he warns against outsourcing thought to AI: “If you get the answer directly, your brain has not changed; your weights have not been updated.” People need to build a “gym for thought.” Education should prioritize agency, not taste—taste may no longer be uniquely human. His closing principle: “Being humbled frequently is a blessing. When someone is never proven wrong, it is probably not because they are always right, but because they have stopped progressing.”
Deep dive
1. Being proven wrong is the norm: the “model king” changed hands 4 times in 6 months
- He opened by owning the miss: in the previous episode, No. 124, he forecast 2026 as the “year of R,” expected a correction in the second half, and exited his entire public-markets book before the Lunar New Year. “People kept sending me my old podcast and asking whether I had been proven wrong.” His response: being proven wrong is the norm in early-stage investing—“When you are repeatedly proven wrong, it means the industry is changing very quickly, which means there are a lot of opportunities.”
- The narrative timeline ran as follows: OpenAI ruled in November and December last year, with DAU heading from 800M toward 1B and Oracle’s stock jumping dozens of percentage points after OpenAI announced its compute purchase; Gemini 3 launched in December, and “Google has TPUs, the most compute, and native multimodality,” making it the new model king; Claude coding exploded in January and everyone began learning agentic coding, while Anthropic’s revenue surged in March and its public-market valuation surpassed OpenAI’s, creating a narrative in which “Anthropic was pulling away from everyone”; in May, Codex’s new-user growth overtook Claude Code, GPT 5.5 performed well, and OpenAI “came roaring back.” There was also an interlude in which “Office 4.7” (original audio) briefly seemed to make the model less intelligent before the issue was confirmed to be a systems problem.
- The conclusion comes first: “Even the question of which model is better would have had a different answer depending on when you asked it over the past 6-8 months. If you are afraid of being proven wrong, you can say nothing and make no judgments forever—and you will certainly never be proven wrong.”
2. Strong opinion weekly held: the investor who sold everything adds back chokepoint exposure
- He initially did not want to record a second episode. The geek CEO “瓦总”西东 (phonetic) persuaded him: “If you think your views have been proven wrong and therefore no longer want to talk about them, then you are truly bound to your views.” Expressing a view and receiving feedback resembles reinforcement learning, and “the market will not say nice things to you. If the market says you are wrong, you are wrong.”
- The methodology is strong opinion weekly held: a strong view must be grounded in a clear set of reasons. “When those reasons change, a truly smart person should change their view.” The world is fundamentally Bayesian; what happens earlier affects the next judgment.
- In public markets, his model is Stanley Druckenmiller, Soros’s former trader—not a fundamental investor who holds for 5 years. After seeing Anthropic usage surge and using Claude Code heavily himself, he added back storage, optical components, CPUs and other chokepoint positions: “Silicon Valley is broadly going long hardware, especially the bottleneck pieces of hardware.”
- But he remains guarded: “My podcast did affect me. I don’t dare add as aggressively as my friends—the more forcefully you express a view publicly, the more your own words constrain you. That is something I need to overcome.”
3. The review: right about the consumer ceiling, wrong about the coding discontinuity
- He was right about OpenAI’s subscription pricing: it is difficult to take ordinary users from $20 a month to several dozen or $100, and the 50M paying users have already filtered out everyone willing to pay for a chatbot. The slowdown in DAU, paying users and ARPU growth has been confirmed over the past 6 months.
- He was also right that advertising and e-commerce would not be easy. “Even ByteDance, as powerful as it is, spent a long time exploring e-commerce advertising on Douyin and TikTok. Facebook and Google took 6-8 years to become advertising money-printing machines. Everyone was too optimistic about putting ads in ChatGPT.”
- What he missed entirely was the speed of Anthropic’s coding-revenue breakout. Claude 4.5 and 4.6 created a shift “from quantitative change to qualitative change”: “Tasks that many people previously could not complete with Anthropic coding can now be completed very well.”
4. Intelligence is discontinuous: 99 degrees still does not produce a steam engine
- Why did the shift go unnoticed? Internet business models—e-commerce, portals and IM—emerged at the end of the last century and then improved incrementally. AI is fundamentally intelligence, and it only creates value after crossing a threshold. Raising AI from a cat’s intelligence to an IQ of 90 would be a major achievement, “but if you are hiring an employee, you surely want an IQ of at least 100.”
- His signature analogy: “You only get a steam engine once the water boils; at 99 degrees, you still don’t—but predicting exactly what temperature counts as boiling is difficult.” One supporting clue: Claude 4.5 was not called 5. The version number suggests even Anthropic did not expect such a large jump.
- The post hoc explanation is not a new architecture or training paradigm, but high-quality coding-user data combined with large-scale agentic RL. “Agentic loops that previously could not run started running,” and AI began completing programming work over longer time horizons and at higher value.
5. Anthropic’s moat is organizational: bottom-up in exploration, top-down in convergence
- He stressed that this is second-hand information—Anthropic is relatively closed to China—but after speaking with people across Silicon Valley in March, one theme came up repeatedly: organizational capability and organizational form. Exploration favors a bottom-up structure. OpenAI initially had 3 directions—robotics, world models and game-playing RL—while language models were only a small team. That freedom was valuable before ChatGPT.
- There is another side to OpenAI: Sam is “a bit like an angel investor,” as many people in Silicon Valley describe him, constantly encouraging small projects. The result is “kill but don’t bury”: GPTs, “Sara” (original audio), the browser and Operator were launched, then “nobody managed them; it felt like they disappeared again.”
- Anthropic is an organization built for convergence. New hires face values interviews; Darry (original audio) sends the entire company an internal-thought memo every 2 weeks; Claude Code shipped dozens of features in a matter of dozens of days. It “kills and buries”—“very much like a Chinese company with strong execution.” In a high-speed race around a converging mainline, the right structure is one that makes a clear bet early and aligns from the top down.
- But he left room for a reversal: if a new paradigm emerges in 1-2 years, the industry will return to exploration. That is the point of his “second R,” Research, from last year: Silicon Valley’s New Labs are betting that the next paradigm will require organizations that are divergent, free and well funded.
6. Coding went from vertical to horizontal, and Claude Code relieved the “API-only” concern
- One major reversal is that, until last year, even senior researchers grouped coding with healthcare and finance as a vertical. Now everyone is realizing that coding is horizontal: it accelerates work in offices, healthcare and research. “Once you see it, it becomes obvious; before that, it was not obvious at all.”
- The second reversal is that 姚顺宇 worried as recently as mid-last year that Anthropic sold only APIs and had no moat—what would happen when a cheaper API arrived? Claude Code answered the question: high-quality usage data flows back from users and trains a better coding model. “Without Claude Code, Anthropic as an API-only company would be a very different company.”
- “Claude Code and Codex are both harnesses.” A powerful model alone does not let you sell an API: users use a product, not an API. Anthropic had both a strong model and the first harness capable of running long-horizon tasks.
- A good harness does not have to come from a model company. Claude Code can connect to other models and is not tightly coupled to one model; OpenClaw was built by 1 person and was briefly one of the hottest harnesses. “Last year everyone said wrappers were unimportant. Now everyone says harnesses are important. The judgment about how important this is keeps changing.”
7. The 3-part valuation framework: underpriced short term, overpriced medium term, $5T-$10T over 10 years
- Short term—3 months—the 2 companies’ private-market prices are each around $900B, while the consensus is that either would list at no less than $2T. “If they listed in the current market mood, I would estimate an immediate 2x; a $3T hype valuation is possible.” With year-end AR expected at $100B, a $1T valuation is only 10x revenue and “cannot be called particularly expensive.” On sentiment, both are actually undervalued.
- Medium term—1-2 years—they remain overvalued. The return problem has not actually been solved; it has only taken another form. “It is possible that in 2 years you will find their valuations have fallen substantially.” Over 10 years, “perhaps both are winners and both remain for the long term; they could be $5T or $10T companies.”
- One observation-bias caveat: the leading players have not established a real gap. Whoever releases the last model looks the most impressive. There is a real late-mover advantage, but “overextending that logic can also lead you astray.”
8. Claude Code vs. Codex: from a model war to a battle over brand, distribution and price
- Brand lock-in and habit matter: “Claude Code works pretty well for me, so I have no motivation to try Codex.” Configured skills and settings written into CLAUDE.md are switching costs, and first-mover brand recognition matters.
- Codex is fighting on price. At the same cost, it is roughly 50% cheaper, while friends using both say it is “actually pretty good” on long time horizon and reasoning. GPT 5.5 changed how people perceive Codex: “A good horse needs a good saddle.”
- By mid-May, in the 5th month, the gap between the 2 products was much smaller than it had been 2 months earlier. The competition has become a battle over brand, distribution and price. Coding plans tie model discounts to the harness, so users need to distinguish between a model that is good because it is discounted and a harness that is intrinsically better.
9. The return problem is merely being pushed down to Anthropic’s customers
- The core argument is that capex ultimately has to become returns. Anthropic’s revenue surge has led many to conclude that the return problem is solved, but Anthropic’s revenue is its customers’ investment. “People are not buying this many tokens to burn them for fun. They ultimately want a result.”
- The chain has 3 steps: input → output → result. Tokens buy software output; that software must be sold or reduce costs and increase profits before it becomes a result. Revenue growth requires new products and new businesses. “Has mobile internet stagnated in recent years because there are not enough programmers to write code, or because people do not know what to build? I think often it is not a shortage of programmers.”
- There is a timing mismatch. Discovering new products and new demand is a gradual process, “but the tokens you burn are burned now.” GPU clusters are designed around payback periods of more than 6 years, but “can the tokens I burn today be monetized 2 years from now? Most companies cannot withstand that.”
10. Tenfold engineers do not mean tenfold revenue; organizations have limited responsibility capacity
- Consider a thought experiment: if a scaled company suddenly had 10x as many engineers, would revenue surge? “Often, it would not.” There is abundant anecdotal evidence of individuals becoming 10x more efficient, but we have not yet observed a major improvement at the outcome level.
- Silicon Valley is beginning to copy one another faster: programmers are cheaper, but no one knows what new thing to build. “The first reaction is to build what someone else has built.” Lovable (pronounced “Labble”) started building PPTs; Gamma started building websites. Replicating existing features is getting faster, but creating features no one has imagined remains hard. “That is not a programming problem. It is an innovation and product problem.”
- The bottleneck for knowledge workers is accountability: an organization has a finite amount of responsibility it can absorb. AI can write 10x as many reports, but people still make the investment decision. “AI cannot be held responsible, and I cannot fire it.” Reports still hallucinate and require human review, bringing the bottleneck back to people. “Being able to review 10 companies does not mean you can invest 10x as much money.”
11. Cost reduction means layoffs, but layoffs have 3 constraints: the sink model
- First, programmers spend 80-90% of their time coding, but another 10-20% goes to communication, coordination and interpersonal trust. “That part cannot be directly replaced by an agent.” Autonomous driving could handle 90% of highway driving early on, but corner cases still prevent removing the human.
- Second and third, large Silicon Valley companies can cut 15% of existing excess headcount at any time, whether or not AI arrives; deeper cuts are much harder. And layoffs are a one-time action: “You can only lay off a person once.” You cannot keep generating incremental cost savings by laying off the same employee. Those laid-off programmers’ income is also revenue for other companies, making the macroeconomic shock substantial.
- The sink model: countless people who want to try the new technology are flowing in from the top, while some users who decide the crab is not worth eating are leaving from the bottom. “Right now, many more people are coming in than leaving.” That is why Anthropic, OpenAI and Chinese coding models all have surging AR: “If you have compute, you can sell it.” But the leak must eventually be sealed: early adopters must actually make money from the tokens, and the cycle cannot be too long.
- His own experience provides evidence: “In March, I burned far more tokens than I do now, because many of those tokens produced no fundamental value.” The beneficiaries are small companies and individuals with ideas who were previously blocked by a lack of people to execute them. Large companies face the problems of “not knowing what to build and having very low organizational efficiency.”
12. Too big to ignore: $700B in hardware profits demands an answer
- Why ask about returns now? The numbers are too large. “If it were tens of billions of dollars, people could afford it. But AI hardware profits across the industry will be $700B this year.” Samsung and SK Hynix are approaching Nvidia’s profit level; workers are receiving annual bonuses worth several million renminbi; and traditional CSP cloud providers are already less profitable than memory manufacturers.
- If Anthropic reaches $100B in year-end AR, that means selling $10B worth of tokens every month. “How much money are the people spending $10B on tokens making? This question increasingly needs an answer quickly, otherwise it becomes too big to ignore.” Hyperscalers are already borrowing to build data centers, Amazon among them, and leverage is shortening the time available to find the answer.
13. The 2026 narrative: agents and the pleasure of being an AI boss
- The main line has never changed: AI raises productivity, and agents are the only meaningful form—“Only agents that do not need human attention can liberate human attention.” He recapped his own thinking: in 23, he discussed humans working closely with multi-agent systems; in 24, he discussed model improvements unlocking value; programming, computer use and reasoning unlocked 25 as the “Year of Agents.”
- His investing is consistent with that view. He backed Manus and Genspark, the first wave of harnesses, and invested in Kimi, which “should be the Chinese model company most focused on agents and coding.” Kimi completed a major transformation in 25 and “may now be the best open-source coding model in the world.”
- One sociological observation: this year, people have become addicted to agents and “always feel there is more work to do.” A friend captured it in one line: “We are actually enjoying the pleasure of being the boss.” You issue an instruction and reality changes. From attention is all you need to everyone becoming the boss of an AI.
14. Anatomy of a harness: the F1 driver and the pit crew
- The metaphor is straightforward. Previously, using a product was like driving yourself; now it is like watching F1. The model is the driver, while we are the crew changing its tires, maintaining it and keeping it on the track. Harness originally means tack for a horse: “making something extremely powerful operate within the boundaries you specify.”
- Outside the model, the layers are context—real-time, organization-specific information that “does not fall from the sky into the model”—then tools and the agent loop, with sandbox and runtime at the outer layer. OpenClaw’s appeal lies here: it gives the model broad permissions to operate Mac files and runs a heartbeat every 30 minutes to check unfinished tasks. “These mechanisms sound simple when explained, but they let the model perform better on long-horizon tasks.”
- Runtime is now gaining traction. Manus worked with E2B last year on a sandbox that let models browse the web and use software; the current direction is to give AI a persistent computer of its own. DigitalOcean’s virtual-machine business has grown sharply. “More and more value is sitting in these layers wrapped around the model.”
15. The model is the product, and the wrapper is being revalued: both statements can be true
- On 杨植麟’s “the model is the product” argument, the core point is that the model provides product capability. But “we always interact with a product, not an API.” ChatGPT itself is a harness: it positioned InstructGPT in a conversational format, and “without that harness, there would have been no such revolution.”
- Cursor proves both statements can hold at once. The market initially thought wrappers had no value, but Composer was trained on Kimi’s pre-training plus Cursor’s large volume of high-quality feedback data. That “shows that even if all you have is the wrapper, the data and post-training it generates can still make ‘the model is the product’ true—without Cursor’s wrapper, there would be no such model.”
- Users are loyal to the harness. People keep “rescuing their own OpenClaw” because they care about the memories stored in the harness, while they switch models casually. After Kimi K2.5 launched, some users replaced Claude with it: “90-point performance, 20-point price.” Memory lives in the wrapper; the model is plug-and-play.
16. “Not fundamental” PPTs: model companies cannot build every harness well
- The disagreement between guests from the prior 2 episodes is worth preserving. 吴力 (original audio, previously rendered as “福利”) believes agent frameworks such as OpenClaw are extremely important and can raise the ceiling of mid-tier models; 姚顺宇 thinks the form factor matters less and focuses more on model-company capability. 戴’s verdict: “They both work on models. If you ask them what matters most, they will inevitably say training a good model—but both actually recognize that a good harness improves model performance and drives data feedback.”
- His signature story: when Manus first appeared, its PPTs looked better than OpenAI’s. He asked an OpenAI agent researcher about it and was told, “That is not fundamental.” “For a model researcher, it is not fundamental. For a user, it is fundamental: if I use you to make a PPT, you need to make a good-looking PPT.” Research ability and application understanding are different capabilities.
- The ecosystem analogy is Microsoft: it built IE and Office, but Windows still supported Adobe and Autodesk. “If applications and harnesses are ultimately all built by model companies themselves, the ecosystem and the use cases will be too narrow.”
17. Harness innovations: what Manus, Claude Code and OpenClaw each invented
- Manus ran a virtual browser inside a sandbox so AI could access websites and execute tasks. It was the first to do wide research: “What we now call agent teams and agent swarms—launching dozens of sub-agents simultaneously—Manus did earlier than Claude Code.” Anthropic also drew heavily on Manus’s experience when building Claude Code.
- Claude Code’s key product decision was the CLI: GUI is for people; CLI is for AI. Skills and configuration of the agentic loop allow the model’s capabilities to be used effectively.
- OpenClaw runs locally on a Mac and can access files, calendars and personal information. Its heartbeat handles scheduled tasks. Most counterintuitively, it puts every conversation into one large context and organizes it into memory.md each day. Researchers dismissed the approach because contexts could bleed together and hallucinations could occur, but users strongly felt that “OpenClaw remembers me.” It also has almost no native TUI usage; instead, it plugs into WeChat, WhatsApp, Discord and Telegram—“go where you are already familiar”—which was key to its popularity in China and is not the kind of innovation a typical model company would pursue.
18. Harnesses are the OS; models are the CPU
- Correcting 广密’s analogy that “the model is the OS”: the harness is more like the operating system, while the model is the processor that runs inside it. In the Windows era, you could plug in an Intel or AMD CPU and “use whichever offered the best price-performance,” just as users swap models inside a harness today.
- For developers, the implication is material. 3 years ago, AI application developers had to write their own agent loops and handle memory and other harness chores. Now Claude Code and OpenClaw handle communication with the model; developers only need to build the skill. It is like Windows freeing developers from handling hardware so they could work through APIs, or iOS encapsulating the camera and processor behind an interface.
- The trend is already visible: applications built on the Claude Code runtime, including skills, Pencil and 真格-backed 斯洛克 (phonetic). There are even graphical interfaces that run on top of Claude Code: “You install Claude Code, but you use an application built on Claude Code.”
19. Do not follow the crowd; be willing to build horizontally
- His first advice to founders: “The biggest fear is following the crowd for safety. You can raise money if you lack cash, and you can try again if innovation fails, but starting with something uninnovative is extremely dangerous in AI.”
- The second contrarian point is to be willing to build horizontally. Everyone in Silicon Valley talks about verticals, but vertical SaaS emerged because SaaS was already mature and the general-purpose opportunities had been exhausted—hence pet-hospital booking systems. Going too vertical early in a technology cycle “can trap you”; horizontal products benefit from changes in the broader wave. Manus and its peers started as general-purpose engines before expanding into PPTs, websites and data analysis.
- The evidence is in the breakout applications of this cycle: Manus, Genspark, OpenClaw and Lovable are almost all horizontal. Lovable is also expanding laterally.
20. AGI is shrinking: within distribution versus genuine innovation
- The hot take is that our definition of AGI keeps shrinking. It began as a singularity capable of destroying humanity, then became solving the Riemann hypothesis and creating new things, and now means replacing ordinary programmers. “刷亚” (phonetic) says AGI has already arrived; he disagrees—that is coding AGI, another narrowing of the definition.
- Writing everyday code is entirely within distribution: “AI writing frontend code is easiest because it is all the same stuff.” Genuine out-of-distribution problems remain open. AI is destroying old value much faster than it is creating new value. New value requires new drugs and new discoveries, which auto research may unlock; many of Silicon Valley’s new New Labs are working on this.
- Even within coding, there are different levels. Distilling existing data into a model is “just copying the answer.” Training models to train models and accelerating the iteration loop is different. “Everyone is doing coding, but some people are doing it differently.”
21. Three defenses against model companies moving down the stack: self-trained models, memory moats and agent network effects
- Model companies moving down into harnesses is “certainly possible.” The first defense is using accumulated data to train a proprietary model—Cursor and Composer are the classic harness-to-model example. “xAI may really want to buy it, but whether it sells is a separate question. Composer’s existence clearly makes the company more valuable.” The second defense is building the moat in the product and memory: users can switch models, but they do not want to switch harnesses.
- The third defense is a major idea he has been developing for the past month: network effects between agents. 6 months ago, your agent and mine were interchangeable, so there was no reason to transact. Now different skills and personalized context have accumulated inside different harnesses, and “the same task will produce different results when handled by different people’s agents.”
- The concrete scenario: “My agent hires 张小珺’s agent to write an interview outline. I give you $10,000: $1,000 is the value of the tokens, and $9,000 is the proprietary knowledge accumulated by your agent. The model may never have knowledge as specific as 张小珺’s.” The current approach—turning the knowledge into a skill and distributing it generously—does not work: “The moment you create the skill, you lose control of it.”
- Companies are already building agent-to-agent marketplaces, though the market is very early. The analogy is 归藏: “The things his agent produces are simply different, and people are willing to pay him to take assignments.” “You can pay a premium to hire 沈南鹏’s agent, or spend a little money to have 戴雨森’s agent review your pitch deck.”
22. The application opportunity is expanding: the capital markets are answering
- The evidence is clear: over the past 12 months, every VC has backed more companies at higher valuations. “A little over a year ago, it was difficult for Manus to raise tens of millions of dollars. Now many application companies emerge at valuations of several hundred million dollars, and ByteDance alone has several of them.”
- The first-principles explanation is that stronger models expand the range of applications. “When models could only write poems and translate, there was not much of an application layer.” As memory, sandboxes and other infrastructure improve, applications become easier to build; the moat shifts back to user data, network effects and brand. “I have friends who are used to Perplexity and will not switch.”
- Large companies are slower at application innovation. Apart from 豆包, “look at what other big companies have produced that makes people gasp—it is actually very little.” US AI applications are also mostly built by startups; OpenAI and Anthropic are themselves startups. “Historically, no company has ever done everything.”
- 真格’s pace is modestly higher than last year but relatively steady; “I know some peers are investing much more.” He personally reviewed 100 companies last year and invested in 2; by May this year, he had already invested in 3.
23. Five traits incumbents dismiss: this time it is wrappers, open source and insecurity
- The framework is that every era’s major companies grew while incumbents watched, because incumbents always had reasons not to act: too niche—Airbnb, Bilibili and 得物; too low-end—Pinduoduo and 内涵段子; too labor-intensive—Meituan’s ground operations; noncompliant—Uber, Didi and Bitcoin; or too far ahead—OpenAI and SpaceX.
- The mapping to today is wrappers—“If you think wrapper value is zero, what if you are wrong? If it is not zero, there is room to add value”; open source—OpenClaw can acquire users through open source and potentially sell a paid enhanced version; and insecurity—OpenClaw reads your files and crashes frequently. Large companies are reluctant to build such products because they would be responsible when something goes wrong; that is precisely the advantage of startups that “move fast and break things.”
24. Invest in thoroughbreds before themes
- In the second half of last year, he made his first investment in 2 embodied-AI companies: 刘松明, born in 2000, author of the RDT series and working on large-scale UMI data training roughly in parallel with Generalist (phonetic); and 丁宁, born in 1997, author of the SimpleVLA series. Both are young Tsinghua PhDs and professors. Their valuation was roughly RMB200M at the time; “now they seem to be worth several billion renminbi.”
- The logic is not reverse-engineering a theme: “We did not decide to invest in world models and then search for founders in that direction—we already believed these 2 thoroughbreds were exceptional. Where they chose to run was for them to decide.” After tracking them for several years, he invested as soon as they started companies, naturally becoming the sole and largest investor in their first round.
- The precedent is 植麟: in 2021, when few people even knew what a large model was, “he was already training a large model.” The traits were deep understanding of the frontier, entrepreneurial spirit and the ability to abandon many things and start decisively.
- The story of the dexterous-hand company “五机” (phonetic) is similar. After graduating from UIUC, the founder used his family’s money to build high-torque motors. He later concluded that the best way for robots to manipulate objects would be a hand, that human-hand data would be easiest to obtain, and that the cost of hands would fall rapidly. He kept building high-degree-of-freedom dexterous hands when one hand cost several hundred thousand renminbi and everyone called it too expensive for too small a market. Now egocentric video data has made dexterous hands structurally similar to human hands strategically important. “It was not that we understood from day one that hands would be especially valuable. An excellent founder explored his way to that conclusion.”
25. The founder framework: 4 types of people, 4 forms of force and original judgment
- This is not pure intuition. Founders fall into 4 categories: prodigies, seasoned operators, scientists and managers. Internally, they often discuss 4 forms of force: learning ability, creativity, leadership and willpower. The underlying capabilities do not change with the era: “It is hard to imagine that one field needs smart people while another does not.” Since 徐小平’s New Oriental days, he has consistently focused on founder chemistry, trust, motivation and whether the team is complementary.
- What changes is the technical dimension: look for “leaders”—people working on agents or world models before everyone else was talking about them. The key test is original judgment. “Many people’s judgments come from podcasts they listened to. Some people can generate original judgments. That may be extremely important.”
- A PhD is not mandatory. Research-driven directions such as world models require actual research experience—刘松明 left school during his second or third year of a PhD. AI applications do not require a PhD, “but your understanding of models may need to be at the same level as a researcher’s.”
26. AI product managers still need personal heroism: Peter and the origin of OpenClaw
- 姚顺宇’s view is that the era of individual heroism for AI researchers is over. Once the training pipeline is industrialized, “the person working on pre-training data is not that different from another person doing the same thing.” But AI product managers still have an opening. 戴 agreed and went further: good products have always required individual heroism. Jobs, 汪涛 (phonetic) and 张小龙 all had “a view of the future that most people did not see, then assembled a team to build it.” That requires a bias about the future.
- The requirement is a deep understanding of the AI frontier: what can be done now and what will probably be possible in the next 6 months. “Good AI products are generally designed for the future.”
- OpenClaw founder Peter is the model. A seasoned operator whose company sold for $100M, he was a power user of Claude Code. It began with the question of how to keep using Claude Code while eating. He built a hook to connect the agent on his computer to an IM service for remote control, and step by step it became OpenClaw. “He had a need no one else had yet—the leading edge of the user base.”
27. The old rules no longer work: attention formulas, walled gardens and subsidized acquisition
- The business-model formula has changed. Mobile internet was DAU × time spent × monetization efficiency: attention is all you need. The constraint was that there are only so many people and each has 24 hours, so companies had to deepen monetization through livestream tipping. In the agent era, the key metric is the task duration that can be sustained, as measured by METR—how long an agent can run after being given a task. Attention is not all you need.
- The moat has reversed. Super-apps once locked users inside closed apps, but closed ecosystems that agents cannot access become liabilities. “If I chat in WeChat, the agent cannot see it; if I chat in Feishu, the agent can see it—and Feishu has even launched a CLI. I now have an incentive to move my work groups to Feishu. Your former moat becomes a barrier restraining yourself.”
- User acquisition has also reversed. The Spring Festival chatbot red-packet war was an old mobile-internet reflex. Leading AI applications spread through magical product experiences—ChatGPT, Sora, Manus, Veo 3 (phonetic) and Seedance 2 (phonetic). 豆包’s highlights were also product magic, such as dialect recognition and its input method, rather than pure paid acquisition. “Users acquired through red packets leave again because there is nothing magical about your experience.”
- The user has changed from a person to an agent. “How many people use it every day is no longer important; how many agents use it is more important.” CLI and API need to be excellent; the GUI does not need to be especially fancy.
28. DAU is not the north star: 10 fanatics beat 100 indifferent users
- Claude Code’s DAU is far below ChatGPT’s, “but its revenue may be much higher.” Optimizing DAU creates absurd outcomes, such as forcing users to return more often just to take another look. A better product is one where “you hand it a task and it runs on its own, so your visits actually decline.” Falling DAU can mean the product is doing its job. The north star for productivity products is long-horizon execution.
- What matters early is the quality of engagement: “100 users who think you are OK may be worth less than 10 users who love you.” Some products will explain that they are the TikTok of the AI era, “but after all the logic, the product simply is not fun.” AI games and the metaverse failed because “they really have no use.” A good product is like a good dish—it is not good because it sounds good when described; it has to taste good.
29. The 3-step agent rollout, and 3 types of software after coding becomes infinitely cheap
- AI is like an alien arriving in the human world, and the transition comes in 3 steps. First, give humans more and better agents—Claude Code, Codex, OpenClaw and Manus are all working on this, while installation, configuration and management become easier. Second, make agents adapt to the human digital world, which was designed for people through GUIs, CAPTCHAs and credit cards. AI used to be a bot being blocked; now Stripe and Coinbase want to issue AI credit cards, while Cloudflare is shifting from blocking AI to allowing agents to register and use services equally. “It is like a foreigner coming to China: you have to arrange a SIM card, WeChat and Alipay.”
- Third is agent-native infrastructure. Human payments are infrequent, point-to-point and large; agent payments are frequent, small and many-to-many. Writing a report and querying 100 databases could require 100 payments. “There may be no concept of a card at all.” AI uses CLI and API rather than GUI; “GUI exists because humans are very limited—we cannot remember where we were looking, while AI does not have that problem.”
- Once coding capacity becomes unlimited, 3 new types of software emerge: everything built by elite programmers—ByteDance’s top engineers build recommendation algorithms, while ordinary engineers build smart kettles, which is why those apps are difficult to use; personalized software—“I use the same WeChat as my grandmother in her 90s, and that itself can be improved,” such as a voice-message progress bar and 2x playback for 徐小平’s frequent 60-second voice messages; and one-off, low-frequency software—“Maybe today, when we record a podcast together, there will be a web app built specifically for it.”
30. The AI-native organization: full-stack teams and the lesson from steam to electric motors
- Companies will get smaller and specialization will blur. The scale required to reach PMF in software will shrink. Waterfall divisions—architecture, frontend, backend, testing, UI and operations—will become full-stack pods: “A few people own a functional module, from foundational design through launch and operations.” Specialization existed because human context is limited; AI does not face the same constraint.
- Day-one context is an advantage. In a company that has operated for 10 years, “the biggest problem is that your context and data are not inside the AI, and moving them in is extremely difficult.” New companies such as 斯洛克 (phonetic) use their own platform for operations, task management and coding—a case of “build Slack using Slack.” Claude Code was built using Claude Code; Codex was built using Codex.
- The history lesson is organizational. Steam-engine factories were built around a central shaft, with long, narrow layouts. Electric motors used wiring to free the factory from that shape; only large, flat factories enabled Ford’s assembly line. “The shift from steam engines to electric motors did not automatically produce productivity gains; it required a change in physical organization.” Big companies cannot deploy AI by merely installing Claude Code for everyone. Organizations with rigid departmental walls will struggle to adapt to AI. Human change may take 10 years—or require new people to replace old ones.
31. The next ByteDance will not look like ByteDance
- He summarizes many AI applications as “AI feeds”: open the app and scroll through AI-generated content in a single or double-column feed, often built by ByteDance alumni. “This wraps new technology in the shell ByteDance is best at, then competes on promotion and distribution—beating yourself under your own game rules is extremely difficult.” OpenClaw is the counterexample: “It does not even have its own home turf; it lives everywhere.”
- Consumer entertainment—kill time—is not the major opportunity of this era. Distribution is controlled by existing champions and newcomers must buy traffic. “You have to be more fun than Douyin, RedNote and Xiaohongshu from day one.” When Kuaishou and ByteDance emerged, they only needed to be more interesting than having nothing to do. The AI-game dilemma is that “the game really is not fun.” Users do not want to use AI; they want to play a fun game. Productivity is the “easy problem”: turning manually built Excel work into AI execution can raise productivity by 10x or 100x immediately.
- To founders from the ByteDance ecosystem: “Their underlying quality is certainly high. My simple observation is that some of them need to go through the process of learning what ByteDance has mastered and then overturning themselves. Building another feed application makes it very difficult, in my view, to find a major opportunity.”
32. The penetration rule: native opportunities emerge after new technology is poured into old bottles
- Every technology revolution starts by using new technology to solve old problems. The internet began with email, portals and self-operated e-commerce; mobile internet began with mobile browsers and mobile search. “But mobile search was still Google and Baidu, and mobile YouTube was still YouTube.”
- The biggest opportunities appear after penetration crosses a threshold and native models become possible. Once everyone was online, social networks, search engines and platform e-commerce emerged. Once smartphones and 4G were widespread, short video, livestreaming and recommendation engines emerged—“the screen was small and people did not have time to search and click, so they needed recommendations”—along with miHoYo’s mobile games, Meituan and Didi’s O2O models, and Pinduoduo in lower-tier markets. “Every company I just mentioned was a startup.”
- Applying the same logic to AI, we are in the first phase, in which AI performs humans’ existing work. Once every person and business has an agent, “perhaps my digital twin will interview your digital twin.” If AI is the one viewing ads, “do ads still exist?” AI-to-AI communication will move data directly through APIs rather than Excel. “Microsoft’s former moat would come under significant pressure” as the value of SaaS interfaces is bypassed.
- Early forms are visible. On 斯洛克, “the 3 of us and 5 agents collaborate in one channel, and you can give my agent commands.” Meta’s acquired Malt Book (phonetic) lets each person’s agent post; “it still feels a bit like cosplay.” But if agents truly run long-horizon tasks for a month, an occasional post would become normal—“like a bar that foreigners regularly visit.”
33. 3 Silicon Valley observations and the China-US contrast
- First, coding has entered a major acceleration. Everyone is using coding aggressively—“use it first, whether it is useful or not.” Meta has a token-burning leaderboard to see who can burn the most. “There is certainly a lot of waste, but some innovation will also be discovered.” Sentiment is polarized: enthusiasts think AGI is near; those who fear unemployment are buying Anthropic private shares. “If you are going to eliminate my job, I might as well become your shareholder. At least I need to get on the train.” It resembles the rush to work for Samsung and SK Hynix in Korea.
- Second, world models are hot in both China and the US. Silicon Valley wants to transfer the language-model paradigm to robotics: massive data and some form of scaling law in exchange for zero-shot and few-shot generalization and stability, corresponding to the 3 stages of pre-training, fine-tuning and IL. The route has not converged: large-scale egocentric first-person video competes with Generalist’s (phonetic) UMI gripper-data approach. The concept itself is contested. 谢赛宁 says “video generation is not a world model,” and that none of these companies is building a world model. “Nobody really knows what a world model is.” The only common ground is that, like a language model predicting the next token, it predicts the next state of the world.
- Third, auto research aims at recursive self-improvement: AI improving AI. 田渊栋’s Recursive has also been announced and appears to be working on self-improving AI. Every top lab in Silicon Valley is using AI to accelerate the model-training process; Chinese research-oriented companies are considering the same.
- The China-US contrast is notable. Silicon Valley now has more than 60 New Labs by some counts, but VCs such as Benchmark avoid expensive, research-driven companies without a clear direction; 70% of YC Demo Day companies are still vertical SaaS, reflecting “habitual momentum.” China is more hardware-heavy: robotics, world models, AI hardware, quantum computing and controlled nuclear fusion. Horizontal general-purpose applications are unusually scarce in Silicon Valley—“Manus does not really have competitors right now.” Perplexity and Devin have reached $400M-plus in AR, “and I believe there are still many opportunities in horizontal applications.”
34. The consumer-electronics hot take: AI hardware will repeat the fate of new consumer brands
- The hot take: “Many consumer-electronics companies today will repeat the fate of the new-consumer wave.” In 2020, new-consumer brands rebuilt every category and used 李佳琦 to sell them; the market later learned that small innovations plus livestream selling did not create a new category. Many AI hardware products simply add a chatbot or some sensing capability, without solving a fundamental problem. Hardware is also harder than software: supply chains, tooling, slow iteration, capital intensity and sales complexity all weigh on the model.
- As an agent vehicle, the market keeps circling back to the Mac mini. Phones and computers are already well balanced across size, performance, power consumption and heat dissipation; a dedicated “OpenClaw machine” is inferior to a phone plus the cloud. Wearable recording devices are, “to put it bluntly, a need invented by VCs.” VCs imagine themselves as busy and needing 10 to-do reminders, but for most ordinary people, the incremental value of recording every day is limited relative to the hassle.
- 真格 has made almost no wearable investments. He wears Oura and Whoop himself: “Those are health hardware, not AI hardware. They need to be genuinely useful, rather than being worth a lot simply because they use AI.” The projects he passes on most are copycats. Plaud (phonetic) is “actually very good because it defined a product form”; it was an “exam setter.” The devices that followed, taking various shapes and plugging into the back of phones, were imitators, as were companies building “a Manus for every industry.”
35. Robotics: invest in components and brains, not humanoids
- His portfolio includes components—五机’s (phonetic) dexterous hand and 方舟’s robotic arm—along with 2 world-model companies. He invested earlier in 非夕 (phonetic) and its spinout 穹彻 (phonetic), but did not invest in this wave of humanoid-robot companies.
- The reason is that humanoids remain “very, very early,” with the work still mainly research. Whatever the narrative, “usefulness depends on manipulation,” which is why he invests in hands, arms and models. “Humanoids are currently in a phase of extreme capital enthusiasm, with valuations high relative to shipments. I do not know whether that can continue. We have certainly missed many opportunities, and we are watching.”
36. Information diets and outsourced thinking: building a gym for thought
- Podcasts are an excellent channel for information distribution, and “people will synchronize with frontier knowledge faster and faster.” But they also mass-produce “judgments heard on podcasts.” He does not listen to podcasts himself. Instead, he built an AI tool that checks transcripts from shows he follows every day, pulls them into Notion and lets him read the text. “A 7-hour podcast does not take me 7 hours to read.” 谢赛宁’s episode was one of the few he listened to in full.
- The bigger problem is outsourcing thought to AI. “Previously, when you encountered a problem, you would think about it yourself. Now you ask AI directly. If you get the answer without going through the thinking process, your brain has not changed and your weights have not been updated.” The analogy is a wheelchair: sit in one constantly and your leg muscles atrophy. Humanity first outsourced physical labor to machines, then memory to the internet, and is now outsourcing thought to AI. Karpathy says you can outsource thinking but not understanding; 戴 goes further: “Thinking may not be fully outsourceable either.”
- The solution is a gym for thought: deliberate thinking the way you deliberately exercise in a gym. “To prepare for this podcast, I forced myself to clarify the arguments behind my views.” Innovation also requires deliberate practice, starting with a simple web app, “just as fitness begins with stretching and push-ups.”
- One model is “林总,” who burns $10,000 a month on tokens. He built a scanner for expiring domains; whatever domain the AI registered became the product he built. “It is like Omakase: you cook whatever fish is available at the market that day.” He also built real-time monitoring for his 5 cats. 戴 built a family health dashboard aggregating Whoop and Oura data, then co-created a health-comparison app with friends on 斯洛克. “The original goal was to make me lose weight.”
37. Unemployment, redistribution and human-AI dependence
- Unemployment “may indeed be unavoidable.” Technology is spreading faster than many people can relearn and adapt. Individuals can protect themselves by learning to use AI or moving into areas driven by human trust and responsibility. “No one can afford the cost of failing to develop.” Redistribution will probably take a form similar to UBI or taxation. In Korea, some have already proposed that Samsung and SK Hynix are earning too much; such measures would also support social stability.
- The long-term optimistic case comes from the diffusion chain of the steam engine: mine drainage, which initially allowed only linear motion; Watt’s condenser and planetary gear; the spinning jenny; more clothing per person; demand for color; synthetic dyes, when purple was once the most expensive color; BASF and Bayer growing from dye workshops into chemical giants; then gasoline, plastics and fertilizer; and higher agricultural productivity. “New jobs will emerge in places you cannot imagine, but diffusion takes time.” The Industrial Revolution took decades to spread; AI has replaced large amounts of programming in 1 year, so the disruption may be much greater.
- The biggest variable is how people relate to AI. “The more you use Claude, the stronger the trust and dependence become, and you cannot help telling it more.” OpenClaw users resemble people raising a child: it remembers them, gets work done and even writes in a relatively human way when organizing files, producing emotional projection. “If you remove AI from your life today, it may genuinely be difficult to remove it. We were not prepared for this at all.”
38. Education: agency over taste, with OOD as humanity’s final territory
- What should children learn? “It is a difficult question.” 2 years ago, he said humans would retain agency and taste; he has now withdrawn taste. “AI is so intelligent and has seen so much high-quality data that its taste is far beyond that of ordinary humans. Taste is not a reliable refuge.” What remains is agency: no matter how good AI is, “you still have to tell it what to do.” People need to ask good questions, have problems they want to solve and want to change the status quo—although proactive agents are challenging even that.
- Humanity’s final territory is OOD. Existing model architectures handle within-distribution problems well, but AI cannot tell an original joke; it can only refresh jokes humans have already told. Galois developed group theory in his 20s. Sutton says humans still have not learned the intelligence of squirrels or cats. “If all your capability lies at the center of the human normal distribution, you are easy to replace.”
- History keeps distilling away humans’ core capabilities. The Industrial Revolution eliminated physical labor, making education a route to mobility. The internet separated knowledge from capability, making knowledge searchable and execution valuable. AI is separating execution from judgment, and taste may also be replaced by AI running experiments against rewards. What will remain important, at minimum, is the agency and drive to decide “what to do last”: “At least for now, AI cannot move itself. I am thinking about this, but I do not have an answer.”
39. Epilogue: being proven wrong frequently is a blessing
- He will keep recording. 瓦总’s line landed: “If you feel you have been proven wrong and then stop recording, you have truly been defeated.” Expression itself is deliberate practice. “When you explain something, you discover which parts you genuinely understand and which parts you have merely glossed over.”
- The closing principle for the entire episode: “As an early-stage investor, being proven wrong frequently is a blessing. When someone is never proven wrong, it is probably not because they are always right, but because they have stopped progressing.” Being proven wrong provides feedback and drives growth; it also signals that the industry is changing, which is where early-stage investing finds opportunity. “Enjoy being proven wrong.”