2026 AI New-Year Conversation: the year of R | A Conversation with ZhenFund's 戴雨森
Summary
- 戴雨森 believes 2025 marked the shift from AI’s “BlackBerry era” into the application-driven “iPhone era.” O1 introduced thinking-time scaling; Claude 3.5 Sonnet brought coding capability; OpenAI o1 added reasoning; and Anthropic enabled long-horizon planning. Together, they pushed GPQA, SWE-bench and other capabilities past the usability threshold, triggering the breakout of Cursor, Claude Code, Codex, Manus and Genspark. The coding-agent market went from zero to potentially around $1B in ARR within a year. “Progress in model capability unlocks application opportunities” remains the conversation’s central thesis.
- The Agent thesis has been validated, but product maturity remains at the early-market stage; renaming a workflow does not make it an Agent. A real Agent derives from agency: independently decomposing goals, selecting and calling tools, and changing course based on feedback. Its core output is “saving people time.” Devin, Manus, Claude Code and 豆包手机助手 are all only beginnings. 戴雨森 agrees this is the “first year of Agents,” while stressing that it will be the “decade of agents”: moving from L2 assistance to L3 responsibility could still take years.
- 戴雨森 defines 2026 as the Year of R, with the first R standing for Return: markets spent the past 3 years trading the I in ROI, and must now validate the return. Tens and hundreds of billions of dollars in data-center investment, $100M annual talent offers, and the rallies in Nvidia, optical components and memory are all driven by the potential return from AGI and application profits. But moving frontier models from 80 to 90 requires far more investment while capability gains are slowing, and Chinese open-source models can reach 80%–90% of SOTA within roughly 6 months. The gap between “high expectations and slow deployment” may become a key variable for market sentiment and financing conditions in 2026.
- All 4 mainstream monetization paths can grow, but none automatically delivers the grand valuations previously assigned to them. Chatbot subscriptions at $20 or $200 face token deflation and free competition; advertising and e-commerce may need years to find native formats and will partly redistribute the existing pie of Google, Meta and ByteDance; replacing a $100K programmer does not mean the model can capture $100K, and may instead rapidly commoditize the task; and enterprise AI, despite companies reaching $100M–$200M ARR, still has to cross the chasm of slow large-enterprise deployment. Investors are therefore shifting from pure growth toward gross margin, retention and cash flow, asking whether $1 of tokens can be value-added by an application into $2.
- The second R is Research: as the marginal return on scaling declines, the next level of capability will require new paradigms, organizations and measurement methods. Ilya describes the current phase as a return from scaling to research, while Demis expects AGI to remain 5–10 years away and require 1 or 2 breakthroughs. Self-play, continual learning and world models are candidate directions; New Labs such as SSI, Thinking Machines Lab and Reflection are seeking alternative paths in lower-KPI environments. Meanwhile, MMLU and SWE-bench are nearing saturation. Gemini 3 Pro’s roughly 78 and GPT-4.5’s low-80s do not reflect the gap users feel in practice: without new benchmarks, it is difficult to know whether training is heading in the right direction.
- The third R is Remember: memory may become an application differentiator and move passive assistants toward proactive agents. Existing memory is still like “retrieval with a large notebook”; the next step is using online learning to form a model belonging to each user and understand preferences rather than merely retrieving exact words. 戴雨森 cites ChatGPT recommending Japan’s Yakushima, a remote destination aligned with his niche preferences, as a personal test case. Once AI can combine long-term memory with real-time context to prepare meeting materials and anticipate needs instead of waiting for a prompt, 戴雨森 believes it could open a “10x opportunity.”
- Chinese open-source models are the structural force Koji believes remains underappreciated, and they have changed the table for application founders. Before DeepSeek R1, founders worried they could not access closed-source SOTA. Ten months later, Chinese open-source models had rapidly narrowed the gap, and DeepSeek Math-V2 reached the IMO-gold-medal level of general models roughly 6 months later. Open source acts like a “nuclear weapon,” destroying the short-term lead of expensive closed models while gaining advantages through low cost, transparency and visibility for talent. Application teams should start globally, keep “turning over cards,” and use industry expertise, proprietary data, distribution or integrated hardware and software to avoid being swallowed directly by the models.
- In an environment of rapid model commoditization, the best teams to back are those that can see technical thresholds 6–12 months early and use human agency to amplify creativity. 戴雨森 emphasizes learning, leadership, innovation, willpower and the resilience that is becoming scarcer in the AI era. He also stresses product thinking and marketing: “The saddest thing is not saying the wrong thing; it is working hard to build something that nobody knows about.” AI can reproduce Picasso from a data distribution, but still struggles to create a new OOD style, an original joke or a company. The strongest near-term combination is therefore not AI creating alone, but humans adding the finishing touch and AI supplying 100x or 1,000x leverage.
Deep dive
1. 2025 Moved from the “BlackBerry Era” into AI’s “iPhone Era”
戴雨森 reiterated last year’s core thesis: ChatGPT was impressive, but not sufficient on its own to trigger an application boom. The actual causal chain is “model capability improves → the usability threshold is crossed → new product forms are unlocked.” In 2025, that premise finally became true.
The inflection point can be traced to OpenAI o1 around October 2024. Thinking-time scaling drove rapid gains on benchmarks such as GPQA and SWE-bench. As he summarized it, within a year, scores measuring PhD-level reasoning and real-world coding tasks went from far below human performance to above 80, approaching full marks and surpassing strong human performance.
Coding was the first area to deliver. Cursor became an essential tool for many programmers, followed by the breakout of Claude Code, Codex and other coding agents. 戴雨森’s estimate was that the field went “from zero to potentially a billion” in one year—roughly $1B in ARR.
Applications are no longer confined to chatbots. The US has Harvey and Legora in legal services, and Sierra and Decagon in customer service; China has also seen experiments in AI-plus-education and AI-plus-industry. 戴雨森 therefore revised his view: “We have left the BlackBerry era and entered AI’s iPhone era.”
2. The “First Year of Agents” Delivered Landmark Products, Not Full Automation in One Year
From Devin to Manus and Claude Code, and then ByteDance’s 豆包手机助手, 戴雨森 believes 2025 produced landmark Agents across multiple fields. The “first year” call therefore met expectations, but there is no reason to interpret it as the end state.
He agrees with Andrej Karpathy’s warning about the time horizon: this will be a “decade of agents.” Agents must call resources on a user’s behalf, complete tasks and take on greater initiative and responsibility, so these products will naturally require a longer evolution than chatbots.
Autonomous driving is his analogy. Moving from L2 assisted driving to L3, where people can genuinely release their attention, took much longer than the market initially expected. Demis offers a more optimistic near-term milestone: within the next year, Agents may be able to complete most actions on a computer, with clearly defined white-collar tasks potentially reaching “70 or 80 points, 80 or 90 points.”
3. The Boundary of an Agent Is Agency, Not a Predefined Workflow
戴雨森 insists on a narrow definition: “The word Agent comes from agency—initiative.” An Agent must independently decompose a goal, plan a route, select tools, execute, read feedback and adjust its approach until completion is confirmed.
The value of that feedback loop is not a technical spectacle but “autonomy, the ability to save people’s time.” Users only truly hand over their attention when AI can change course as the environment shifts or exceptions and corner cases emerge.
Many products called Agents merely connect a model to a fixed workflow. They work when a problem follows the preset steps, then have no idea what to do when conditions change. “That is not an Agent,” which also means the number of genuinely capable Agent products today is far smaller than the market label suggests.
4. Agents Remain in the Early Market; Models and Products Must Cross the Chasm Together
戴雨森 borrows the “crossing the chasm” framework: today’s users are mainly innovators and early adopters, willing to try new things and tolerate failure. To reach the mainstream market, model reliability, responsibility boundaries and product form all still need to improve.
The PMF doubts that surfaced repeatedly a year ago are fading, because the industry has sequentially validated 3 questions: do people use the product, do they pay, and can some products make money? Positive examples are now emerging for user value, commercial value and margins.
These examples still depend on specific model breakthroughs. Claude 3.5 Sonnet unlocked AI coding and Cursor; OpenAI o1’s reasoning and Anthropic’s long-horizon planning made Devin, Manus and Genspark possible. 戴雨森 estimates that tool use has currently unlocked roughly 10%–20% of its potential; at around 80%, the range of tasks that can be completed will expand sharply.
5. Unified Multimodality Could Produce the Next Cursor, Not Just a Better Generate Button
戴雨森 places a second clear through-line on native multimodality: text, images, video and audio should no longer be handled by separate specialist models, but understood and generated together within a unified model. “Humans are multimodal native.”
GPT-4o created the “Ghibli moment.” Nano Banana and Nano Banana Pro pushed instruction following and the expression of information inside images into a new phase, turning infographics from design artifacts into high-density information media. “A picture is worth a thousand words” here is not an aesthetic slogan but a statement about information-transfer capacity.
Sora 2 and Veo 3, through better instruction following, fidelity and direct audio-video generation, opened up video and AI-produced manga-style drama. Koji notes that realistic short-form drama will take more time, but comic-style content is already emerging in which AI produces and humans consume.
戴雨森 mentions a visual-understanding benchmark called ZeroBench. Out of 100 points, current models can roughly score only 5; he expects advanced models to reach 60, 70 or even 80 within a year. Elon Musk’s description of humans as “pixel machines” receiving visual input points to a market opportunity: the multimodal-era Cursor or Claude Code has yet to appear, and most capabilities remain trapped inside chatbots.
6. AI’s “Phase Change” Comes from Continuous Warming Crossing a Product Threshold
戴雨森’s metaphor is: “Before water reaches 100 degrees, you can only make coffee; once it reaches 100 degrees, it immediately unlocks the steam engine.” Underlying capability can rise continuously, while the tasks users can complete change in steps after a threshold is crossed.
AutoGPT was already trying in 2023 to make GPT-3.5 and GPT-4 think, execute, reflect and use tools on their own. But “the water had not boiled,” so it could only demonstrate a concept. Once the capability range reached Claude 3.5, especially Claude 3.7, products such as Devin and Manus actually began to work.
The phase change cannot be attributed to models alone. Users consume complete products, not benchmarks. Once a model reaches the threshold, AI-native interaction, task orchestration and result presentation are still needed to convert latent capability into a change ordinary users can feel.
7. Typeless Turns Voice from Dictation into an Input Layer That Understands the User
The 2 AI products 戴雨森 uses every day are ChatGPT, which has accumulated extensive personal memory, and Typeless, an early-stage company he has invested in. On the surface, Typeless is a voice-input method; its ambition is to rebuild natural interaction between people and computers.
Users can speak by pressing any designated key. Typeless removes filler words such as “um” and “ah,” understands “I will make 3 points next” and automatically formats the result as a list instead of mechanically transcribing a string of spoken language.
More important is context and learning. The same sentence can take on a different tone when entered into WeChat or Lark, and Typeless remembers details such as whether the user likes periods at the end of sentences. “Once the words you say have been presented by it, they become what you actually needed.” What Siri and Alexa lacked was not a voice interface, but a model capable of understanding intent.
8. 豆包手机助手 Is the First Prototype That Makes “AI Controlling a Phone” Look Workable
戴雨森 calls 豆包手机助手 a “product prototype” or technology preview, not a mature consumer product. It is still too early, has attracted considerable controversy and is not suitable for ordinary users to rely on directly.
But for the first time, it completed an end-to-end task such as ordering food delivery with reasonable completeness: understanding an ambiguous request, acting across applications and handling several corner cases. The prototype “opened a window,” showing what AI taking over everyday software operations could look like a few years from now.
He admits that the excitement is not as strong as it was at the end of last year. Coding plus Agents had nearly covered 2 core human actions in the virtual world—writing code and using software. More fields are now “building steam engines at full speed” after completing the zero-to-one stage, moving from one to ten rather than creating another wave of the same magnitude.
9. Sunday Advances Embodied AI through Product Taste, Data Infrastructure and Demos
戴雨森 says he particularly likes Sunday, the home robot, and for the first time genuinely wanted to buy a humanoid robot for his home. Its aesthetic design helps it stand out from a mass of similar products, while Tony Zhao and Chi Cheng bring research experience in ALOHA and 5-finger grippers.
The team has also released a data glove priced at roughly $200, with a structure corresponding to a 3-finger robotic hand. 戴雨森 values not the accessory itself but the integrated thinking across low-cost data collection, robot morphology and hardware-software integration.
In some household environments, Sunday uses zero-shot learning to load dishwashers and clear tables, and can even pick up 2 wine glasses at once without breaking them. Both guests believe these demos show a technical approach and practical execution distinct from others, rather than merely creating a visual spectacle.
戴雨森 defends the importance of “being able to make a demo.” From AlphaGo and Manus to Steve Jobs’ product launches and the Wright brothers’ first flight, technological progress has always existed, but a strong demonstration can use media leverage to let talent, capital and the public see its significance early. “A good demo is a time machine”: it can cut through 10 or 20 years and make people today willing to work for the future.
10. 王兴’s Early Appeal Came from Product Intuition and a Grand Narrative Reinforcing Each Other
Koji and 戴雨森 were both heavy users of Xiaonei and Fanfou in college. 戴雨森’s Xiaonei ID was 3527, and user numbers began at 1000, making him the 2,527th user. They could not articulate “user experience” at the time, but intuitively sensed that the products were smooth and that the posters and copy had distinctive aesthetic quality.
Koji recalls that 王兴 combined bottom-up user thinking, constantly asking about specific pain points, with a top-down strategic view. He once compared Xiaonei and Fanfou to “building the capillaries of human information”: the internet had previously built highways, and the next step was to let everyone get on the road at any time and connect with one another.
What Koji took from that experience was not the pursuit of famous people, but “trying to participate in the most impressive thing you can sense, and moving closer to the team you think is the best.” There was no methodology at the time; it was closer to intuition and action triggered by a good product.
11. An Email to an Intern Already Contained 王兴’s Leadership Framework
In 2007, roughly 27-year-old 王兴 opened his welcome email to Koji’s internship at Hainet with: “Our team has an ambitious plan. We are doing something meaningful and challenging: building the best Chinese real-name social network and using technology to improve the lives of hundreds of millions of people.”
What followed was not a job description but 3 alignment questions: what is the company’s goal, what is the individual’s goal, and how can the 2 be coordinated as much as possible? “Only after clarifying these 3 points can cooperation have a solid foundation rather than being a momentary impulse.”
Faced with internship options at Microsoft and Google, 王兴 offered a direct comparison: “If you want to get things done, make a real contribution and gain something tangible, Hainet may be a very suitable place… We have to improve every day.” Koji believes the message still applies to anyone considering a startup today.
When Koji became frustrated by “copying Facebook and Twitter,” 王兴 pointed to the asphalt roads and streetlights around Wudaokou: they were not Chinese inventions either. What mattered was whether users needed them and whether they could be built better and more cheaply. 戴雨森 added: “Learning is not the problem; failing to learn well is the problem.”
12. Bug Reports, Public Expression and a 10-Second Compile Time Revealed Early Talent
When Koji first approached 王兴, he did not offer a generic expression of admiration. He brought a BlackBerry phone and reported a bug in Fanfou’s web version, which happened to hit exactly what 王兴 cared about. 王兴 had also come to know him through his public writing on Fanfou, showing that “building in public” was already a talent-discovery mechanism in the early social-media era.
That generation wrote blogs, posted on Fanfou and commented on Google’s new products. Many of their judgments look immature today, but those activities helped them meet a group of lifelong friends. Koji’s advice is to keep expressing yourself: post on blogs, Twitter, Weibo or Xiaohongshu. Public thinking can become the entry point for collaboration.
张一鸣’s first impression on 戴雨森 during the Hainet period was “an extremely reliable engineer”: fast, inquisitive and capable of putting pressure on product managers in every meeting. An even more telling detail was that while others drank water and zoned out during a 10-second compile, he would mutter, “What could I use these 10 seconds for?” and try to optimize the waiting time away.
戴雨森 adds a Picasso analogy involving “turpentine.” Critics discuss form, structure and meaning, while the real artist may only talk about where to buy cheap turpentine. Founders need both the grand narrative and the turpentine; “keeping your feet on the ground while looking up at the stars” must coexist in the same person or team.
13. The Founder’s “Dao” Does Not Change with the Technology Cycle
戴雨森 reduces founder capabilities that endure across eras to 4 traits: learning, leadership, innovation and willpower. Whether building McDonald’s, an internet company or an AI company, founders cannot know everything at the outset and cannot complete organized creation alone.
Learning handles the unknowns that keep appearing. Leadership brings together excellent people who share the same values. Innovation means that even when borrowing an external approach, the result fits local users and the team’s own conditions. Willpower sustains long-term persistence and resists short-term temptations.
He does not believe AI will produce a completely opposite talent standard. The “technique” will upgrade, while the “Dao” will largely remain the same. The real question is how the next 王兴 or 张一鸣 will express these old qualities under new technological conditions.
14. Chinese Startups Are Moving from Replication and Local Models to Global First Launches, but Competition Is Objectively Harder
戴雨森 divides the past 20 years into 3 phases. The early period was copy to China, such as Facebook and Xiaonei. Around 2015 came models specific to the Chinese market, including bike sharing, shared power banks, Xiaohongshu and Pinduoduo. By 2025, more teams were targeting the global market from day 1 and even becoming the world’s first innovators in their categories.
He views Manus as the world’s first general-purpose Agent, with users potentially coming simultaneously from the US, Brazil and South Korea. Typeless’s voice-input use case is also inherently cross-lingual. Large language models lower language barriers, while knowledge workers around the world perform highly similar tasks, so “China first, overseas later” is no longer the default path.
Starting a company is also harder. Twenty years ago, China’s GDP was growing at roughly 8%, the internet was expanding from a niche into a user base of more than 1B, and the world still offered the distribution dividend of 8B people. Today growth is slower and incumbents are stronger: “Now it is all 张一鸣 competing with you.”
Founders themselves have also improved. Early on, simply studying overseas products early could create an edge; today teams watch launches, study products and learn fundraising in real time with the rest of the world. Capital markets are more mature too. AI resembles semiconductors more than mobile internet: it depends on research, generational iteration, capex and scaling laws, not merely a new distribution channel.
15. When There Is No Map, Speed of Action and Talent Density Are More Reliable Than Forecast Precision
戴雨森 compares 2009 mobile internet with today: everyone sensed an opportunity, but nobody knew the final shape. In an environment without a map or GPS, direction can be based on guesses and assumptions, but action must be fast.
He believes 2 things are almost never wrong: act proactively first, then go where excellent talent clusters. Technology waves often begin in very small geographic or social nodes. Rather than searching alone for a perfect answer, enter an environment with dense feedback.
Concreteness remains a screening standard. Sam Walton’s autobiography describes how the founder would stop in every town during road trips to inspect supermarket shelves and prices. The equivalent slogan for AI founders is “Get your hands dirty”: use ChatGPT, AI coding and Manus personally instead of listening only to secondhand concepts.
16. The CEO, the Number Two and the Investor Are Fundamentally Different Lifestyles
戴雨森 has been the number 2 at a startup, an executive at a listed company and a number 1. He describes the number 1 role as “the most ambiguous and the loneliest,” but also the one with the greatest freedom to create and exercise agency. The number 2 is more like a producer working with a director, responsible for coordinating complex matters and getting them executed.
戴雨森 has long considered himself a “number-two personality.” The number 1 is the last line of defense and may have more power, recognition and returns, but often has lower happiness. If someone does not crave managing thousands of people or pursuing maximum impact, finding the number 1 who fits them best may be the more honest choice.
Koji sees career choice as an investment. Choosing a company and deciding whom to work with means betting nonrenewable time, at a higher cost than investing money. By that measure, he “invested in 王兴,” while 戴雨森 “invested in 陈欧”; early-career returns went far beyond equity.
戴雨森 distinguishes 2 kinds of anxiety. Founders work deeply on one thing with a clear launch timeline, so the work is hard. Early investors know about many topics but lack depth in all of them, and cannot stipulate “produce a unicorn in 3 months”; they are anxious because fate is not fully in their hands. He jokes that investors are like “large models emitting tokens”: they sound as if they understand everything, but may not truly understand anything.
17. Resilience Is the Founder Quality AI Will Amplify First
戴雨森 believes resilience now ranks higher. AI tools let teams build more products faster, which also means more frequent failure. Models can also “drown applications”: the small castle built yesterday may become obsolete after the next release.
The Manus team is his recurring example. After its AI browser ran into obstacles, it quickly searched for a new opportunity, went all in on Manus and built an Agent. The point is not never making a wrong judgment, but experimenting quickly, acknowledging failure without defensiveness, changing direction and restoring team morale.
2 other qualities remain essential: product thinking and marketing. The faster technology changes, the easier it is for people to “carry a hammer looking for nails”; teams must step back and ask whether they are solving a real need. Fragmented attention also means a good product can disappear without a trace, so marketing is no longer merely post-success packaging.
18. The Best AI Startup Entry Point Is 6–12 Months Ahead of Technical Maturity
戴雨森 offers a specific window. Building on fully mature technology is often too late; relying on technology that will not mature for 3–5 years may turn a company into an early casualty. A pioneer should build the product that a technical breakthrough 6–12 months from now will be able to support.
杨植麟’s path came from technical judgment. He anticipated long context and the memory it would unlock in 2023, then shifted early toward agentic systems, making Kimi’s agentic capability a priority. The Manus team, having studied AI browsers and tool use first, anticipated the capability threshold that Claude 3.7 would reach several months later when it began development in October 2024.
AI deals a dozen or 20 new cards every year, allowing teams to experiment repeatedly. Manus was the Butterfly Effect team’s third product, and Typeless was also its team’s third. Max AI had previously competed with Monica; later, one side shifted toward Manus and the other toward Typeless. Staying at the table matters as much as being sensitive to the new cards.
Long-term moats cannot simply overlap with foundation-model capabilities. 戴雨森 values model intelligence combined with industry expertise, proprietary data, distribution channels or integrated hardware and software. Legal, customer service, education, manufacturing and China’s strength in consumer electronics are all areas foundation-model companies cannot cover in one stroke.
19. Being Seen and Taking Small Steps Turn Ambiguous Opportunities into Feedback
Koji’s first media recommendation to founders is “begin with the end in mind.” Appearing on a podcast is not about waiting until everything is ready to collect applause; it requires a clear objective—whether to reach VCs, talent or industry influence—and answers to “Why should people care? Why you? Why now?”
The other half of “know yourself and know the other side” is understanding that media needs good content. Koji compresses it into 3 words: interesting, useful and resonant. A flat feature tour is usually insufficient; viewpoints, conflict, concrete stories and significance are what enter public discussion.
His “hot take” on marketing is: “The saddest thing is not marketing yourself and saying or doing the wrong thing. It is working hard to build something that nobody knows about, discusses or cares about.” Without attention, it is impossible to validate whether the product is right or wrong, and impossible to attract talent.
The first step during an ambiguous period does not have to be quitting one’s job. The developer of Plan Coach, frustrated that he kept failing to wash the dishes, followed AI’s advice to “stand up first” and realized that overcoming procrastination depends on executable small actions. After turning it into an app, a Xiaohongshu post received roughly 260,000 likes. Koji wants founders to view contacting peers and expressing ideas publicly in the same way—as small steps.
20. Return Will Pull the AI Narrative Back from Investment Scale toward Economic Delivery
戴雨森 first “puts on armor” around his forecast: “Everything I say is wrong.” He then defines 2026 as the Year of R, with the first R standing for Return, because the market had previously traded mainly the Investment in ROI.
The $10B data-center projects, hundreds of billions of dollars in plans, Meta’s offer of $100M per person per year and the rallies in Nvidia, optical components and memory are not driven by love of investment itself. They reflect a belief that AGI or AI applications will ultimately generate enormous returns.
The problem is that the experiential gains from new SOTA models such as OpenAI o3, GPT-4.5 and GPT-5 have begun to narrow relative to models from 6 months earlier. Moving a model from 80 to 90 requires more compute and human labor, while frontier progress is no longer as visibly dramatic as it was in the early period.
Large investment has not prevented low-cost catch-up. Chinese open-source models often reach 80%–90% of SOTA roughly 6 months after its release. After OpenAI and Gemini used general-purpose models to win IMO gold, DeepSeek Math-V2 reached a similar level through an open-source approach roughly half a year later, challenging the narrative that simply spending more would keep competitors away.
21. Chatbot Subscriptions and Advertising Will Not Monetize at AGI-Narrative Speed
The first path is subscription. ChatGPT can charge $20 or $200 per month, but 戴雨森 rejects the linear idea that intelligence improves every year, so the price rises from $20 to $200 and then to $2,000. Token prices keep falling; a capability sold for $200 last year may soon be worth only $20.
Netflix is his reference point: its subscription price has not changed by an order of magnitude in 20 years. Chatbot penetration among knowledge workers is already meaningful, and many users are satisfied with the free or $20 tier. Gemini’s capabilities, low pricing and Google’s financial strength will continue to constrain price increases.
The second path is advertising and e-commerce. ChatGPT has surpassed 500M DAU and may move toward 1B DAU. Historically, Google, Meta, ByteDance and Tencent monetized through these 2 paths, but a substantial portion would simply reallocate existing advertising and e-commerce budgets rather than create new GDP.
Native monetization takes time. Google launched in 1998 but did not establish AdWords and AdSense until roughly 2003–2004. Facebook launched in 2005 but did not have feed advertising until 2012. Douyin developed from 2016–2017 and matured only around 2020. Simply inserting ads into a chatbot could also damage trust in an assistant, especially a paid assistant.
22. AI Coding May First Deflate the Value of Wages, While Enterprise Services Still Have to Cross the Chasm
The market often takes the world’s tens of millions to 100M programmers, each earning around $100K, and derives a multi-trillion-dollar AI revenue pool. 戴雨森’s rebuttal is: “AI replacing programmers’ work does not mean it can earn those programmers’ wages.”
Once a model makes a task easy, the task itself is no longer worth what it was. Only the most advanced, SOTA tasks can temporarily support usage-based pricing. As capability becomes commoditized, the task moves into a fixed subscription, then toward free, and may eventually run on-device.
Enterprise services are the fourth path. Companies such as Harvey have reached roughly $100M or $200M ARR this year, but from a very low starting base. Office Copilot has broad use cases and mature technology, yet actual enterprise adoption remains below expectations, showing that large-company deployment and adoption do not happen automatically when a model launches.
戴雨森 summarizes this with Amara’s Law: people tend to overestimate technology’s short-term impact and underestimate its long-term impact. The 2026 issue is not that AI lacks value; it is that the capital already invested demands rapid returns while enterprise deployment remains in the process of crossing the chasm.
23. Genuine AGI Returns Should Expand the Productivity Frontier, Not Redistribute the Advertising Pie
戴雨森 cites Satya Nadella’s standard: if something is to be called AGI, GDP growth should visibly accelerate, potentially reaching 10% annual global GDP growth. “Making the pie bigger” is closer to the economic meaning of AGI than scoring highly on a model leaderboard.
Under this framework, ChatGPT taking advertising from Google, Meta and ByteDance is merely redistribution of the existing stock. AI creating new drugs, discovering new knowledge and raising overall productivity would represent incremental value from expanding humanity’s frontier.
This is why caution about Return does not reject the long-term technology trend. The question is how quickly and how much new return short-term valuations imply, while commercialization may first pass through price wars, trust issues, organizational deployment and deflation.
24. Investors Are Shifting from Pure Growth toward Gross Margin, Retention and Cash Flow
Early in the year, many applications pursued growth at negative gross margin. Cursor was once described as “selling foundation-model tokens at a discount.” The assumption was that tokens would rapidly become cheaper, so acquiring users mattered more than current gross margin.
The new question is whether applications have value-add capability. Even if the input cost of tokens is $10 today and falls to $1 in the future, can the product turn that $1 into $2 that users are willing to pay, rather than continue reselling below cost?
Silicon Valley investors are therefore paying more attention to revenue quality: whether gross margin is positive or negative, whether retention keeps users and whether cash flow is healthy. 戴雨森 mentions that Lovable is reportedly cash-flow positive; Midjourney has never raised funding. These examples are raising the market’s bar for “quality growth.”
A Return shortfall could also change the pace of capital deployment. 戴雨森’s risk combination is “high expectations and slow deployment”: as long as returns continue to meet expectations, investment can continue; once they diverge, market sentiment and follow-on investment will both be affected.
25. Research Is the Second R; the Next Capability Layer Awaits a Paradigm Breakthrough
Ilya’s framework is that AI alternates between scaling and research. Research finds a new paradigm, after which compute and data are scaled; when scaling approaches a bottleneck, the system must search for another breakthrough. 戴雨森 believes the industry is now returning to research.
Dario also acknowledges that AI capability can keep increasing while the speed of economic returns slows, or that economic returns remain uncertain. Demis expects AGI to remain 5–10 years away and require 1 or 2 breakthroughs.
Self-play, continual learning and world models are the frontier directions mentioned in the conversation. The first 2 aim to let models continue updating through interaction with themselves; the third attempts to build an understanding of the world through vision and logic rather than predicting tokens solely from static corpora.
Koji retained another research clue from Ilya’s interview: emotion is not a burden on intelligence but resembles a value function in machine learning. A person who has lost emotion might spend an hour unable to decide which socks to wear; if AI is to improve decision efficiency, emotional mechanisms may deserve renewed study.
26. New Labs Are Seeking the Next OpenAI-Style Breakthrough in Looser Organizations
Silicon Valley is investing heavily in research-oriented New Labs, including Ilya’s SSI, Mira Murati’s Thinking Machines Lab, Reflection AI and Hume.
戴雨森 highlights that Thinking Machines Lab’s seed valuation is reportedly around $50B, even higher than the combined valuation of Chinese generative-AI startups. This shows that capital still holds extremely high expectations for the next breakthrough in research paradigms.
Existing leading model companies already have clear product lines: ChatGPT targets consumers, Anthropic focuses on coding and enterprise services, and Gemini emphasizes multimodality. The New Labs thesis is that genuinely different research paths require new organizations, a looser environment and fewer product KPIs and time pressures.
The controversy cannot be ignored. Research outcomes are highly uncertain, and VC capital and return timelines may not fit foundational research. 戴雨森 does not claim New Labs will necessarily succeed; he treats them as a new US investment trend and believes similar organizations may emerge in China.
27. Old Benchmarks Are Nearly Saturated; Research Needs New “Gaokao Questions”
From early MMLU to SWE-bench, later proposed by 姚顺雨, benchmarks once provided clear training directions. Now many leaderboards are at 80, 85 or 90, with each additional point requiring enormous investment and becoming increasingly disconnected from user experience.
戴雨森 gives an example: Gemini 3 Pro scores roughly 78 in coding, while GPT-4.5 is just above 80. The leaderboard gap is only 2 or 3 points, yet users feel GPT-4.5 is much better. The difference is clearly not just 1%–2%.
Benchmarks are not merely marketing scoreboards. They provide feedback on whether the training, pretraining and post-training path is correct. That is why 姚顺雨’s proposition holds: “AI’s second half has arrived. We need new benchmarks.”
28. Remember Upgrades Long-Term Memory into a Personal Model and Proactive Service
戴雨森 completes the third R with Remember because memory is already changing retention and user experience. He cannot verify whether ChatGPT’s “smile curve” comes from memory or model upgrades, but his own experience is clear: 3 years of history makes ChatGPT understand him better than Gemini.
When 戴雨森 asked where to go for a week during the Spring Festival, ChatGPT did not provide a generic popular-destination list. Based on his preference for remote, niche travel, it suggested Yakushima in Japan—a place he knew and wanted to visit but had forgotten.
Existing memory is still mainly retrieval, like AI carrying a “large notebook” and flipping through old records before answering. Genuine understanding would mean forming a “model of you” internally: even without recalling the exact words, it could anticipate your choices and feedback.
Online learning could eventually give each person a model of their own. Combined with a deep understanding of the user and their context, this would produce a proactive agent. A good assistant does not wait for the boss to call; it prepares meeting materials in advance, reads the room and solves needs proactively. 戴雨森 sees this as a potential “10x opportunity.”
29. 2025 Was Also the Year AI Truly Entered Everyday Life
戴雨森 points to scale signals: 豆包 surpassed 100M DAU, ChatGPT surpassed 500M, and both showed a “smile curve” in which early users returned after an initial period of churn. AI is no longer merely a new tool for industry professionals; it is beginning to form a daily habit.
A friend of 戴雨森’s looked up in a large Hong Kong-style restaurant in Guangdong and noticed that children at 5 or 6 nearby tables were all holding phones and talking to 豆包. While researching an AI presentation for first-grade students, 戴雨森 found that “100%” of 6-year-olds knew about AI and had interacted with it in different ways.
The 2 guests also wrote a short piece for ZhenFund about 2036, while admitting that a 10-year forecast is “basically waiting to be proven wrong.” The purpose is not precise prediction but clarifying what kind of future they want, then deciding which company to start, invest in or join.
戴雨森’s honest reaction to the change over 3 years is: “If someone had told me 3 years ago that AI could do these things today, I would never have believed them.” From the trough of China’s “AI Four Little Dragons” and Silicon Valley’s sense that only 140-character Twitter remained, he believes the era once again permits founders to dream big and shoot for the moon.
30. Chinese Open Source, Human OOD Creativity and the Real World Define the Next Opportunity Boundary
Koji believes the most underestimated force remains Chinese open-source models. Early in the year, founders were still complaining that they could not access model companies’ SOTA and that the competition was unfair. Ten months later, the gap between open and closed source had narrowed rapidly. Open source is not only cheaper and more transparent; it also makes researchers’ contributions visible globally and helps attract top talent.
His metaphor is that “open source is like a nuclear weapon.” A closed-source company may spend $500M training a leading capability, only to see open source catch up 6 months later and destroy the short-term technology rent. Voice, video, Agents, memory-driven proactive AI and hardware-software integration in consumer electronics still offer abundant niches.
For the year’s products, 戴雨森 ranks Manus first. The recent surprise was Tuny AI from a Chinese team. He used his daughter’s name to generate a Christmas song and music video; she replayed it repeatedly, singing and dancing. Compared with Suno, Tuny AI places greater emphasis on 30-second, 60-second and full-length short-video distribution. 戴雨森 says Suno currently has roughly $200M in ARR and has already built powerful user mindshare.
The 2 guests ultimately locate human value in OOD creation. AI can reproduce Picasso’s style from a large body of his work, but struggles to create the next Picasso, a genuinely new joke or an original company. Marc Andreessen’s “founders who use AI to amplify human creativity” are therefore especially scarce: humans provide agency and the finishing touch, while AI adds 100x or 1,000x execution leverage.
31. After AI Raises Virtual Productivity, Happiness Will Depend More on Teams, Bodies and the Real World
戴雨森 gives his own happiness a 9, higher than the roughly 6.5 often cited by many top founders. Koji gives himself an 8 and says the number would have been lower last year. 戴雨森’s missing point is mainly about exercising more and taking better care of his body and family, not dissatisfaction with his career direction.
He believes the growing pains of entrepreneurship are difficult to eliminate, so entrepreneurship is first a lifestyle choice: “Do what you like, live in the way you like, and you will be willing to endure the pain.” Investment happiness comes from good teams, interesting projects, family and the ability to help founders with capital, resources or emotional support without bearing the full burden of entrepreneurship alone.
戴雨森’s top recommendation of the year is A Brief History of Intelligence, which moves from single-celled organisms in the ocean to GPT-4 and examines the emergence of intelligence across 5 stages. A line left by 穿越平行宇宙 is: “It is not the universe that gives life meaning; it is life that gives meaning to the universe.”
Koji uses Alice Munro’s short stories and 远东冰原上的猫头鹰 to step away from AI for a while. 戴雨森 once bought a model of a “plesiosaur transport vehicle” that does not exist in reality. Their shared conclusion is that as AI completes more tasks for people in the virtual world, the giant owl on the ice field, absurd toys and the pleasures of concrete life will become more precious.