127: AI Mid-Year Check-In with ZhenFund's 戴雨森: OpenAI & Kimi K2
Summary
OpenAI’s unreleased general-purpose large language model solved five of the six problems at the 2025 IMO, reaching gold-medal-level performance and marking the episode’s most important capability jump. It had no internet access, no math-specific optimization, and did not use Code Interpreter; despite Google’s objection that the result was not officially certified and 陶哲轩’s reminder that proof scoring can vary, it still suggests LLMs are beginning to crack tasks that are hard to produce and hard to verify. The episode relayed one researcher describing it as a “moon landing moment for AI”: if AGI was once a train smoking in the distance, “we can hear it now.”
The breakthrough strengthens three technical theses at once: inference scaling, general-purpose generalization, and scientific discovery. More thinking time can continue to improve performance, while the model’s base model was reportedly the same as GPT-4o, implying that much of the gain may have come from post-training and inference; if true, substantial room for optimization remains even if pre-training hits a wall. IMO proofs and unproved minor theorems are structurally similar, leading 戴雨森 to argue that AI may be close to discovering “small new knowledge,” with potentially greater productivity value than coding tasks that mainly copy, adapt, and assemble existing code.
The value of both model capability and the application-layer “shell” is being underestimated. In the same weekend, the model delivered gold-medal-level IMO performance, while ChatGPT Agent’s actual outputs, including PPTs, lagged behind some online results from Manus, Genspark, Kimi, and MiniMax; model progress does not automatically eliminate applications. Agents need applications to provide organizational and personal context, cross-session memory, tools, and execution environments, and “the same model understands me better” could become a durable moat.
By the first half of 2025, coding and reasoning had crossed the chasm, while Agents entered the early-adopter phase of mass adoption. Cursor, Claude Code, o3, Deep Research, and Kimi Researcher have demonstrated clear productivity value, while Manus and Genspark are beginning to handle goal decomposition, tool selection, execution, and review. True L3 is not a chat window with a few extra buttons; it is “humans stop doing the work, AI does the work,” with users shifting from operating tools themselves to learning how to be AI bosses.
Kimi K2 is a case study in the market underestimating a strong team. As of the recording date, July 15, 戴雨森 called K2 “the best open-source model in the world, period,” with particular enthusiasm for its coding, agentic workflow, and Chinese writing; its OpenRouter coding call volume climbed from No. 13 to No. 10 within days, offering a more direct user vote than benchmarks. Behind it are a stable core team, a long-standing bet on long context, renewed investment in pre-training, and continued exploration of RL, tool use, and open source; K2 was still a non-reasoning model, with reasoning and multimodal versions yet to come.
Agent-driven compute demand could dwarf the Chatbot era, and Nvidia’s move past $4T may not be the end of the road. 戴雨森 observed that Agent applications can consume more than 1,000x the tokens of ordinary Chatbots, just as dial-up-era assumptions that everyone would only chat on QQ could not have forecast the bandwidth required for 4K video. Productivity demand is also not directly bounded by individual leisure time: AI can cover 50 stocks in parallel and process multiple earnings calls at once. In his view, the conversion of tokens into productivity has only just begun.
The binding constraints have shifted from whether AI can land to talent, organization, safety, and the distribution of gains. Meta’s “disruptive” compensation has reset the industry’s cost base, but whether assembling a large group of star researchers creates a coherent team remains unproven; embodied AI is the opposite case, with lower Optimus production expectations underscoring that manipulation and productization cannot be rushed. By year-end, the key tests will be whether Agents can raise one-shot delivery success from roughly 20% to 70%-80%, whether memory can generate true compounding returns, and how society governs AI code no one can fully review, content that is difficult to distinguish from reality, and widening gaps between individuals.
Deep dive
1. A General-Purpose LLM Pushes the IMO to Gold-Medal Level
OpenAI said its unreleased model solved five of the six problems at the 2025 IMO. It had no internet access, no math-specific optimization, and did not use external tools such as Code Interpreter. OpenAI also asked 3 IMO gold medalists to cross-check the solutions and concluded that the proofs were correct.
戴雨森 stressed that this is not the same achievement as Google’s AlphaProof and AlphaGeometry reaching silver-medal level last July. Those were specialized systems built for mathematics and formal verification; this was described as “a general-purpose large language model,” making generalization the first-order significance.
The model answered soon after the problems were released, reducing the likelihood that it had seen them in advance or been trained specifically for them. Google nevertheless noted that the result had not been officially certified by the IMO, while 陶哲轩 pointed out that different proof paths can produce different scores. The episode preserved those caveats rather than equating “gold-medal level” directly with an official gold medal.
2. Hard-to-Verify Breakthroughs Are Closer to the Real World Than Coding
Go has an unambiguous win-loss result, and code is relatively easy to verify through execution and test cases, giving reinforcement learning a clear reward signal. IMO proofs are hard to produce and hard to verify; determining whether they are correct is itself much more complicated.
曼祺 cited Jason Wei’s point about the “asymmetry of verification”: for some tasks, verification can take longer than generation, meaning that even if AI can do the work, humans may not be able to fully exit the loop. 戴雨森’s response was that producing a proof at all is already a huge technical and intellectual advance; verification difficulty does not erase the significance of that step.
That gives the math breakthrough greater extrapolative significance than coding. Much entry-level code is essentially copying and modifying existing modules, while mathematical proof requires constructing more complex new reasoning. 戴雨森 believes AI is following humans from simple tasks toward tasks that are harder and more valuable.
3. Inference Scaling Opens New Growth After Pre-Training
OpenAI said that giving the model more thinking time continued to improve its score, further supporting an inference-scaling law. 戴雨森 divides model progress into pre-training, post-training, and inference: even if pre-training is said to be “hitting a wall,” the latter two stages may still have a long way to run.
The episode mentioned an unverified rumor that the system behind the result may have used the same base model as GPT-4o, with the real difference coming more from post-training and inference. If true, the result was not simply a larger base model; the post-training phase had reopened the capability frontier.
曼祺 noted that OpenAI researchers also mentioned multi-agent systems. 戴雨森 speculated that “agent” here was more likely multiple internal roles collaborating to strengthen reasoning than a product-level Agent calling different tools. He cautioned that the event was recent and that the details should await disclosure and reproduction.
4. The IMO Makes “Discovering Small New Knowledge” Feel Near-Term
戴雨森’s reasoning is that an unfamiliar IMO proof problem and a small unproved theorem are structurally similar: both require constructing an argument without an existing sequence of steps. AI will not solve the Goldbach conjecture anytime soon, but helping discover “small theorems and small proofs” may no longer be merely a future possibility.
This is also why he sees the math breakthrough as having greater productivity spillovers than coding. Theoretical disciplines that advance primarily through reasoning rather than physical experiments could be accelerated first. The episode ended by asking whether the first PhD students and researchers to become fluent with AI will be the first to turn this capability into genuine scientific discovery.
The episode compared AGI to a train: “We used to see smoke in the distance; now we can hear it.” The comparison still has limits—an IMO is not a complete research problem, much less months or years of autonomous research—but it brings AI one step closer to becoming an “AI PhD student.”
5. In Two and a Half Years, the “AGI Spark” Became a “Moon Landing Moment”
The March 2023 paper Sparks of AGI was excited by a GPT-4 preview model solving problems that were astonishing at the time. Just two and a half years later, those examples already look extremely simple. 戴雨森 noted that two and a half years can be shorter than the product-development cycle of a venture-backed company, yet the intelligence curve was extraordinarily steep over that period.
One researcher involved called the IMO result “a moon landing moment for AI”: a model that superficially only predicts the next token and has no tools produced creative proofs that only a tiny number of human geniuses could complete. The episode argued that the importance of the event “may be impossible to overstate.”
The episode’s early-year prediction of a “Lee Sedol moment” is spreading from Go into coding and mathematics. The point is not that AI has reached ordinary human performance, but that it is approaching or surpassing top human performance in an increasing number of fields. If Google is hinting at comparable results, the implication would be that the technology is no longer held by a single lab—although whether those results come from a general-purpose model remains unclear.
6. Once the Barrier Breaks, Capability Diffusion Usually Follows
After speaking with top Chinese researchers, 戴雨森 found that their reaction was not “we expected this,” but widespread shock. That, in his view, means a new research direction has been made explicit. He compared it to the atomic bomb: “Once you know an atomic bomb can explode, actually building one is not very far away.”
His expectation is that researchers will gradually understand how the result was achieved and improve their own models. The IMO score may therefore represent not only a point lead for OpenAI, but a broad uplift in reasoning capability across the industry.
Uncertainty remains around verification. If a proof requires a small number of experts to review it for a long time, AI cannot yet independently replace researchers. But the ability to generate the proof first will still change the division of labor between production and verification; a human will remain responsible for the final judgment.
7. The Talent War Proves Talent Is Valuable—and Creates Organizational Shock
Meta’s disruptive compensation packages for top researchers are not only rattling the teams being raided; they are also destabilizing Meta’s existing AI organization. Those who stay will ask whether they should receive raises and how their prior contributions are being valued.
戴雨森 acknowledged that the experience and know-how of top talent are extremely valuable, but raised a question obscured by compensation headlines: the better the people, the more coordination they require. Putting many star researchers in one organization does not automatically create a strong team. History offers plenty of examples of prestigious teams failing to collaborate.
China is accelerating its own talent raids. Based on his limited observations, 戴雨森 said Tencent is recruiting from Tongyi and ByteDance, taking over the baton from ByteDance’s previous phase. He explicitly said he lacked a complete industry-wide view, but could confirm that the competitive temperature is rising.
For startups, the opportunity cost of talent, token subsidies, and funding needs are all rising at once. For large companies, exchanging cash for time is a relatively straightforward problem. The talent war therefore validates the value of the sector while continuing to raise the barrier to entry for startups.
8. ChatGPT Agent Was Underwhelming, but Made the Agent Direction Industry Consensus
The first reaction to ChatGPT Agent was disappointment, especially with its PPT output, which critics called “really kind of ugly.” 戴雨森 still sees it as an important milestone: from Devin’s prototype, through early products such as Manus and Genspark, to formal integration into ChatGPT, goal understanding, planning, coding, tool use, and reflection have moved from concepts into a product form backed by leading companies.
Structurally, it combines Operator, which leans toward “writing” operations, with Deep Research, which leans toward “reading” operations. Some people summarized it as “OpenAI launched a Manus,” but the episode said that describes the product form, not a complete comparison of capabilities and resources. Products such as Manus also make their process more visible, helping users understand what the Agent is doing.
OpenAI classified the product at a high-risk level in its AI safety capability framework, citing concerns including phishing sites, biological-weapons information, and other unauthorized actions, and imposed many restrictions. 戴雨森 sees the caution as responsible, but also as evidence that larger companies act more conservatively, leaving startups more room to move quickly, experiment, and push boundaries.
9. Chinese Agents Show That Applications Will Not Automatically Be Outplayed by Model Makers
Manus, Genspark, Kimi, MiniMax, and others subsequently ran OpenAI’s showcased tasks with their own Agents and delivered better online results in areas such as PPT generation. 戴雨森 first saw this as a continuation of Chinese product strength: the experience accumulated in mobile-internet products such as Doubao, JiMeng, and Jianying does not disappear when the underlying technology becomes a large model.
The previous assumption was that once a model company entered an application category itself, end-to-end training would let it “crush” applications calling an API. That logic was closer to reality in the Chatbot phase, when users mainly talked directly to the model and applications had limited room to add context and tools.
The Agent phase is different. Models need more context, more tools, and more complex environments. Application companies must also use context engineering, product design, and long-term practice to provide inputs and execution conditions that the model itself does not have.
The contrast over the same weekend produced a two-sided conclusion: the IMO shows that the pace of model evolution is underestimated, while ChatGPT Agent’s output gap shows that application value is underestimated. If both models and applications are being underestimated, AI as a whole may still be undervalued.
10. Context Engineering Is Turning from a Technique into an Application Moat
戴雨森 compares prompt engineering to a boss giving orders, while context engineering is closer to “context, not control.” It is not merely about writing a more specific brief, but about providing examples, background, data, resources, and permissions in forms the model can use more effectively.
The first layer is context quality within a single session: whether the data is complete and formatted for action. The second is cross-session memory and personalization, allowing today’s, tomorrow’s, and the following day’s work to accumulate in the same assistant. Having better context “with me” could make the same model irreplaceable in the user’s experience.
The third layer comes from new data channels created by products. Glasses that continuously observe the surrounding environment can provide real-time visual context that internet giants did not previously possess. A model cannot obtain that information from nowhere; hardware and software products must supply it.
Manus’s shared context-engineering practices were well received because they address a problem that is not yet standardized and requires long-term trial and error. An application’s value lies not only in choosing a model, but also in how it provides, organizes, and uses context across sessions.
11. User Data Does Not Directly Increase Intelligence, but It Can Continuously Improve the Work Experience
戴雨森 agrees with the part of the argument that chat data cannot create a flywheel of model intelligence. Models have already surpassed the average person, and ordinary users’ preferences on questions with standard answers are unlikely to provide the same high-quality intellectual gain they once did under early RLHF. “If I can already solve IMO problems, why should I care whether ordinary people prefer this answer or that answer?”
Agent task data is different. What the user wants to accomplish, how the workflow unfolds, what gets sent back for revision, and whether the final delivery meets the goal can all improve how the product completes tasks, even if they do not make the base model smarter.
The distinction is therefore between two flywheels: model intelligence and product experience. Ordinary chat may not strengthen ground-truth capability, while organizational processes, personal preferences, and execution outcomes can enrich context. The earlier a company builds an Agent and accumulates these relationship data, the more its value may compound with time rather than with model versions.
12. Coding and Reasoning Crossed the Chasm in the First Half of 2025
戴雨森 heard that OpenAI has divided the business into GPT, API, and coding lines, reflecting the fact that coding has become “the most useful, the easiest to pay for, and extremely fast-growing” direction. Cursor and Claude Code do more than autocomplete; they handle larger codebases, self-correct, and take on more complete engineering tasks.
Reasoning progressed from o1 in the second half of 2024, to DeepSeek R1 at the start of the year, and then to o3 in April. Deep Research and o3 also helped drive ChatGPT user growth. Kimi Researcher was one of the earlier broadly available Deep Research products in China and received positive feedback.
戴雨森 believes these two capabilities are no longer impressive research demos, but have “officially crossed the chasm into the mainstream market.” Users are willing to pay for clear productivity value in coding, and commercialization across other products is gradually being validated.
13. Agents and Multimodality Have Moved from Toys to the Early-Adopter Market
Devin first showed the outline of an L3-level Agent. Manus, Genspark, Claude Code, and others then showed that given a vague goal, AI can find tools, choose a path, evaluate results, and deliver. They remain error-prone and unstable, but have become useful among early adopters and are beginning to attract usage and payment.
The key change in GPT-4o image generation is not simply better “sampling,” but a clear improvement in semantics and instruction following. Ghibli-style images, comics, flowcharts, avatars, and public-account cover images have become practical use cases, and images can now serve directly as output components for Agents.
Veo 3’s addition of synchronized sound made 戴雨森 feel for the first time that video had “crossed the uncanny valley” and entered a virtual world difficult to distinguish from reality. Multimodality is therefore no longer just entertainment content; it also raises the ceiling for expression and delivery in productivity applications.
Financing and project competition have also heated up with commercialization. China’s early-stage market has again seen fights over projects and multiple rounds of financing in a single month, while Silicon Valley has seen large startup financings, acquisition battles, and talent premiums at the same time. The heat reflects revenue and usage beginning to materialize.
14. Digital Productivity Is Arriving, While the Timeline for Embodied AI Was Overestimated
They are currently more focused on AI Agents and productivity tools in the digital world because those use cases can already create direct value. Embodied AI is hot, but Tesla’s lower production expectations for Optimus reinforce 戴雨森’s view since 2023 that the market has underestimated the difficulty of manipulation.
The expectation that 10,000 robots would enter factories to tighten screws in 2025 was too aggressive. Even as demos improve, having a robot reliably make a cup of coffee remains extremely difficult. A breakthrough could still happen this year, but breakthrough, scale-up, productization, and mass rollout are a continuous process and cannot be skipped.
The food-delivery price war may divert resources from Alibaba, Meituan, JD.com, and others. 戴雨森 said bluntly, “Rather than free milk tea, they should be training advanced models.” But Chinese knowledge workers use workflows similar to Americans’ Office, search, and research tools; as Kimi, ByteDance, and DeepSeek catch up, inference demand and cloud capacity will likely replicate the US growth pattern.
The resumption of H20 sales to China and growth expectations for Alibaba Cloud are peripheral signals that demand may be picking up. They have no direct causal link to the food-delivery war; the core point remains that once models are good enough, application adoption will quickly translate into inference-compute demand.
15. Kimi K2 Executes a Comeback Under Pressure Through User Feedback
As of the July 15 recording date, 戴雨森 called Kimi K2 “the best open-source model in the world, period,” emphasizing coding, agentic workflow, and Chinese writing. He also preserved the comparison boundaries: R1 had been out for roughly half a year, while Grok 4’s strengths were not limited to these areas.
Benchmarks have numbed many users, so the episode cared more about real calls. In OpenRouter’s coding category, K2 climbed from No. 13 to No. 10 within days, behind mainly closed-source models such as Claude, Gemini, and GPT. DeepSeek V3 ranked No. 8, in part because it launched earlier and had more existing application integrations.
戴雨森 expects K2’s usage to keep changing as actual users compare capability, cost, and open-source status. Pablosky’s founder also said publicly that he was trying to call K2 and that Kimi was doing very well. These developer choices are closer to a real product vote than a single leaderboard result.
OpenAI had originally planned to release an open-weight model in July before delaying it, creating a dramatic contrast with K2’s launch. The episode did not accept the joke that K2 caused the delay; it treated the episode as an example of the model race being a constant game of catch-up in which no lead can be extrapolated statically.
16. Kimi’s Core Assets Are a Stable Team, Technical Conviction, and Long-Term Accumulation
戴雨森 sees Kimi as different from some “little dragons” among model companies that have undergone major changes in co-founders or business heads. Its core team remained stable under pressure. Many members were classmates, roommates, or bandmates at Tsinghua, giving them a working foundation that predated the creation of a large-model company.
The team built its technical appeal around 杨植麟, with a mission long described as “exploring the optimal solution for converting energy into intelligence.” It did not repeatedly change direction with trends such as companionship, multimodality, or 2G; it stayed focused on intelligence and productivity.
Kimi’s 2023 bet on long text and the launch of a search-enabled Kimi were early evidence of the team’s technical vision. The original use case was putting The Three-Body Problem into context; today, Agents need to execute hundreds of steps, coding requires handling huge codebases, and reasoning requires generating long chains of thought. The importance of long context has only increased.
People did leave during the external downturn, but core researchers stayed for more than compensation. They stayed because they could “learn things” and “actually build something impressive.” 戴雨森 defines K2 as a stage-by-stage answer delivered jointly by technical vision and organizational resilience.
17. K2’s Technical Catch-Up Started with Pre-Training, Not Just RL
DeepSeek R1 once triggered the discussion that pre-training was unimportant and post-training mattered more. 戴雨森 explicitly rejects that simplification: without the strong V3 base model, R1 would not have had its subsequent “left foot stepping on the right foot” reinforcement-learning gains.
Kimi’s RL team was already strong, but its base-model capability had been weaker. Kimi Researcher was still built on K1.5, using RL to improve deep-research capability. The key change over the past several months was concentrating resources on pre-training and producing K2, with 1T total parameters and roughly 32B active parameters.
K2 is still a non-reasoning model, with reasoning and multimodal versions yet to be released, so it is more like a new foundation. Some users currently consider its inference speed slow, but 戴雨森 views that as an engineering issue to optimize rather than a verdict on capability.
The team advanced optimizations including Muon, MuonUp, and MuonClip in its training pipeline and adopted architectural ideas such as MLA. 戴雨森 considers mutual borrowing normal; real capability comes from building the pipeline oneself and solving stability problems, which gives the team greater control over the model.
18. The Value of Open Source Is Control and Community Feedback, Not Running a 1T Model on a Laptop
In response to the criticism that almost nobody can deploy a 1T-parameter model locally, 戴雨森 pointed out that a full-size DeepSeek R1 is also not easy to run on an ordinary computer. Open source does not mean single-machine deployment is mandatory; APIs such as OpenRouter can support broad use of open models.
More importantly, developers can inspect and control the model, preserving options for privacy, private deployment, post-training, and customization. K2’s modifiability and agentic capability give developers an incentive to embed it into real workflows.
DeepSeek demonstrated the payoff from open exchange: when a team gives the global community a genuinely useful model, users will test, spread, fix, and improve it on their own. 戴雨森’s qualification is clear: “Open source is not a magic cure-all.” The prerequisite is still that “you open-source something good.”
19. DeepSeek’s Pace Shows the Need to Choose a Breakthrough Wedge When Resources Are Limited
戴雨森 has only a peripheral view of why R2 has been delayed. DeepSeek may still be training V4; following the sequence in which V3 came after R1, R2 may also be waiting for a new base model. The team faces training-compute constraints, and after going viral it also has to devote substantial resources to inference.
DeepSeek was once more determined to focus on large language models, then increased investment in multimodality, showing that its intelligence roadmap is also evolving. When resources cannot match competitors across the board, it must choose the direction that matters most and is most likely to produce results, as Anthropic did, then catch up elsewhere through engineering and research.
ByteDance has more resources, so it separates responsibilities across Edge, Flash, and Base: some teams explore the frontier, some chase SOTA, and some serve applications. 戴雨森 approves of the division because the long-term goals of model research can easily collide with an app’s bugs, corner cases, and short-term demands.
Kimi has also reflected on this conflict over the past six months, reducing application maintenance and promotion while directing more effort toward the next-generation model and RL. The principle the episode distilled was: “Get one wheel working first.” Application companies should freely adopt the best model for each task, while model companies should first defend a specific SOTA position.
20. The Main Theme in Reasoning and Coding Is Moving from Good to Excellent
戴雨森 believes most usable progress in 2025 has not come from a new paradigm appearing from nowhere, but from quality moving from one to ten. o3 is a step above o1, while o3 Pro uses longer inference time to reduce detail hallucinations and improve rigor; even smaller models such as o4-mini can score highly on reasoning.
Public benchmarks increasingly resemble a high-school exam with a maximum score of 150: a model scoring 150 and another scoring 140 do not necessarily differ by only 10 points in real capability. As leaderboards saturate, internal evals, complex real-world tasks, and long-term user experience become more important.
Coding progress continued from Sonnet 3.5 to Sonnet 3.7 and Claude 4. Context is longer and self-correction is better; in Claude Code, when facing large codebases and complex tasks, users are beginning to experience the model getting it right in one pass.
Alibaba Cloud CTO 周靖人 does not even regard the o1 series as a completely new paradigm. 戴雨森 agrees that Transformer may still have a long way to run, just as highways and railways can continue to carry growth; there does not necessarily need to be something radically different every year.
21. Tool Use Is Moving from an Application Add-On to a Native Model Capability
Agent capability depends on three lines of progress: reasoning, coding, and tool use. Recently, models such as K2 and Grok 4 have begun emphasizing the inclusion of tool-use data during training itself. Native tool competence can reduce application call costs and make end-to-end reinforcement-learning optimization easier.
There are broadly two tool-use paths. MCP is similar to an API, allowing a model to call capabilities through standard interfaces. The other uses visual understanding and interface operation to simulate how humans use browsers, virtual machines, and existing software. Operator and Manus’s sandboxes exemplify the second path.
Software tasks often have clear rewards, such as whether an order for a cup of milk tea succeeds or whether code passes its tests, making them particularly suitable for reinforcement learning. The application layer still has to decide which public and private MCPs to connect, what permissions to grant, and how to write the result back into the real environment.
22. Video Is Approaching Reality, Voice Has Crossed the Threshold, but Bandwidth Remains Limited
Veo 3 became a representative video SOTA model by synchronizing image and sound. MiniMax’s Hailuo 02, Midjourney Video, and ByteDance’s related products are also competing rapidly. Viral examples of kittens, puppies, and elephants diving show that models can generate more complex and continuous visual events.
戴雨森 is particularly bullish on Midjourney Video’s quality and low cost, while also recognizing MiniMax’s and ByteDance’s progress in the vertical. Video remains a model war requiring heavy training and heavy compute, driven mainly by well-resourced large companies and a small number of leading startups.
Voice is also crossing the line into “indistinguishable from reality,” but 戴雨森 believes spoken communication has limited bandwidth. The amount of information it can carry per unit of time is lower than visual media, code, and complex task execution. Voice is an important opportunity, but not necessarily larger than reasoning, coding, and tool use.
23. Google Has Returned from Disappointment to the Top Three in Model Competition
Gemini 2.5’s reputation, GCP serving, TPU compute, and Veo 3 together represent a marginal improvement for Google. 戴雨森 believes its technical accumulation, talent density, funding, and proprietary compute are all deep, making it one of the 3 most competitive companies at the top of the field.
The market remains concerned about AI’s impact on the search homepage and legacy businesses, keeping Google’s stock volatile. This is the classic situation in which the old business is damaged, the new business is catching up quickly, and the net result is still unclear. 戴雨森 thinks the answer will only become clear in another 1-2 years.
Google has a broad strategy spanning multimodality, VLA, and robotics. Some capabilities have moved from good to better, while others remain at the zero-to-one stage. The episode also mentioned hearing that teams such as Gmail are working extremely hard, challenging the old image of Google as a place to coast.
The 3 leading labs are developing distinct centers of gravity. OpenAI increasingly resembles an application company, choosing to pursue 1B users and keep expanding ChatGPT; Anthropic is driven by coding and Claude Code; Google relies on coordination across its full-stack models, cloud, TPUs, search, and multimodality.
24. Application Value Comes from Three Layers: Model, Context, and Environment
The first layer is the model capability fixed in the weights, which sets the basic ceiling for reasoning, coding, and knowledge. But once the weights are trained, they are “dead”: the model does not know today’s weather, the company’s new files, or the decision the user just made.
The second layer, context, includes public, organizational, and personal information. Search fills in public facts; company processes, documents, and knowledge bases provide organizational context; personal conversations, preferences, and history determine whether an assistant truly understands its user. The accuracy, timeliness, and usability of that context can create enormous differences in experience.
The third layer is the environment: which tools AI can call, which systems it can operate, and which real-world outcomes it can affect determine whether it is a question-answering machine or a worker. Application companies provide MCPs, browsers, codebases, documents, and permission boundaries, ultimately turning model output into deliverables.
戴雨森 uses a car analogy to explain the limits of making only the model: the engine is obviously critical, but a complete car later needs “a refrigerator, a television, and a sofa.” OpenAI, Google, and Anthropic will naturally build first-party applications because selling only the model may not showcase 100% of its capability. Whether DeepSeek goes this far depends on its own choice.
25. Heavy Trial Use Leads 戴雨森 to Spend Nearly $1,000 a Month on AI
戴雨森 still uses ChatGPT most often and considers o3 the most balanced overall. In China’s network environment and for his research work, Kimi Researcher is a good fit. When he needs PPT structure and a finished product, he uses Genspark; when he wants to turn the result into a website or interactive display, he tries Manus.
He buys the roughly $200 top-tier plan for Manus, Genspark, ChatGPT, Gemini, Grok, and others, bringing his monthly spending close to $1,000. His rationale is not to collect tools, but that “you have to try new products more,” because seeing future capabilities can feed back into product and investment ideas.
曼祺 said this was equivalent to hiring a person. 戴雨森 replied that an analyst usually costs more. If AI can write more code, complete research, or generate revenue directly, paying by usage becomes natural: “You make me $100, and I pay you $1 or $10.”
26. Agents Push Token Consumption from Chat Scale to Work Scale
After Manus launched in March, 戴雨森 saw inference usage surge. An Agent searches, reads and writes files, calls tools, and repeatedly checks its work; Agent applications can consume more than 1,000x the tokens of an ordinary Chatbot. This became a key basis for his view that inference-compute demand is being underestimated.
He compared his earlier skepticism about Nvidia with a miscalculation from the dial-up era. If everyone were assumed to use the internet only to chat on QQ, enormous bandwidth would not be needed; then 4K video arrived. Estimating AI compute from simple Q&A similarly misses the new use cases that appear as models become more capable.
OpenRouter cannot provide an absolute measure of the entire market because large products such as Manus and Genspark call official model APIs directly. It is better suited to observing relative votes from small and midsize developers and which models are moving quickly up the rankings.
After Nvidia crossed $4T, the episode discussed whether it could reach $5T. 戴雨森 responded that “the trend of converting tokens into productivity has only just begun.” Once the train is moving and keeps finding new uses, demand does not automatically stop after a single efficiency improvement.
27. Productivity Demand Expands More Easily Than Attention Demand
Companionship and entertainment consume a finite amount of time each day, while productivity can scale in parallel. With Cursor and Claude Code, people do not merely write the same code faster; they write more code. With Deep Research, they also ask questions they would never have attempted in the past.
戴雨森 gave a concrete public-markets example. During US earnings season, 5 or 6 SaaS companies may report simultaneously at 4 a.m., followed by conference calls 2 hours later. A human analyst cannot cover them all, but within the next 6-12 months AI may track 50 stocks in parallel, update models, listen to calls, check pre-set questions, and generate reports.
His analogy is that before airplanes existed, nobody had a demand to fly from China to the US for a business trip; demand is not expressed when the capability does not exist. Once AI lowers delivery costs, a large amount of research and analysis previously abandoned because of time and staffing constraints will emerge.
曼祺 challenged 戴雨森 as someone who “really wants to work.” 戴雨森 replied that users will pay for productivity if it generates income. The two retained different emphases on the relationship between productivity demand and personal time.
28. There Is No Single Agent Answer Yet; Model-Native and Application-Orchestrated Systems Will Coexist
戴雨森 defines an Agent’s job as accepting a large goal, choosing tools, planning a path, evaluating progress during execution, and ultimately delivering a result. For now, humans still set most goals, while the model’s ability to decompose and proactively define objectives remains early-stage.
In OpenAI’s five-stage framework, Chatbot maps to ChatGPT, Reasoner to the o-series, and Agent begins at stage 3. The industry has not yet settled on what a truly Agent-native model looks like. Kimi is exploring a “model-level Agent” by adding end-to-end RL and tool trajectories during training.
Manus does not train a base model. Its path is to understand the capabilities future models will have, then unlock them through context engineering, product design, and tool environments. Its “less structure, more intelligence” philosophy emphasizes letting the model explore, although the episode acknowledged that completely unstructured systems are not always optimal.
Genspark adds more explicit tools and best practices for slides, improving output controllability and reducing cost. MetaStone’s Deep Research uses a different interaction model. Having been “proven wrong too many times,” 戴雨森 refuses to name a winner in advance and believes the same task can have multiple effective solutions.
29. The Key to L3 Is Not Better Assistance, but AI Taking the Wheel
The autonomous-driving analogy for L3 is that the human can temporarily stop touching the steering wheel. In work, the core idea is the same: “humans stop doing the work, AI does the work.” Cursor keeps programmers in the editor, while Claude Code goes further: users mainly describe the goal, inspect the result, and issue the next instruction.
Manus observed that when product managers use Cursor, they do not look at the intermediate code at all; they look only at the chat panel on the right because the code means nothing to them. If an Agent product still requires users to understand and edit every step themselves, it will struggle to expand its user base.
戴雨森 describes the new skill as “learning to be the AI’s boss”: provide goals, resources, context, and authorization, then let the AI employee decide how to execute. When users fail for the first time, they often think, “I might as well do it myself,” just as managers instinctively complete a subordinate’s work. But that control prevents the system from accumulating capability.
This is the management interpretation of “context, not control.” In the past, only a small number of people had the opportunity to develop employees; now almost everyone will need to learn how to develop tools. Whether individuals can establish good feedback, memory, and work habits may become a new source of productivity divergence.
30. By Year-End, the Real Test Will Be Success Rate, Memory, and Society’s Capacity to Absorb the Change
General-purpose Agents will ultimately converge on a small number of first-principles use cases. 戴雨森 uses the steam engine as an analogy: it was first used to pump water from mines, and only later found railways and textile mills. The clearer Agent destinations today include coding, PPTs, and Deep Research; future ones may include earnings calls and continuous office workflows.
Investment mindsets differ between China and overseas. China often begins by asking, “Won’t a large company definitely build this?” Overseas investors get excited by the possibility that a startup could challenge OpenAI or Google’s super-app ambitions. After Manus launched, people said it could be copied in “one weekend,” but months later there were still few products with a comparable experience. First-mover reputation, distribution, and accumulated context all have value.
The Windsurf-related transaction showed the intensity of competition. OpenAI was reportedly considering a roughly $3B acquisition, which later fell through, before Google cut in; another initial proposal was reported at $2.4B and primarily involved the founder, several dozen core employees, and investors, while Google did not acquire the remaining company’s equity. The episode treated the details largely as industry gossip. These acquihires give some founders an exit, but may cause ordinary employees’ option returns to diverge from those of founders and core teams, while forcing companies such as Cursor to raise more money for talent and tokens.
The test 戴雨森 most wants to see is whether, under the assumption that current success rates are roughly 20%, Agents can move from occasionally delivering a score in the 70s or 80s to achieving 70%-80% reliability—letting users state a requirement and have the Agent “grunt away at it, then get it right the first time.” Good products must be designed for models several months ahead: Cursor did not take off until Sonnet 3.5 and 3.7, just as YouTube positioned itself for broadband and short video for smartphones and 4G.
The other major question is memory: can AI become like a long-term employee, understanding the user better with every collaboration and becoming harder to replace? That requires online learning as well as continuous access by the application to files, feedback, and personal context. Neither the model nor the product can be missing.
A more distant but rapidly approaching issue is evaluation and governance. GPQA scores have risen from the single digits to above 80, while Humanity’s Last Exam has climbed from a few points early in the year to above 20. When multiple models are smarter than the evaluators, internal evals themselves will become a competitive capability for teams.
The host and 戴雨森 both discussed the real-world consequences. One person may use “cyber farm labor” to build a unicorn, and large organizations may manage more complex businesses, but the gaps between people and between organizations will widen. Vast amounts of AI-written code no one can fully review, bugs, backdoors, privacy leaks, and content difficult to distinguish from reality are already pressing problems.
戴雨森 ended by expressing fatigue with sensational headlines and hollow marketing. Models once seemed like mysterious wizardry; now the industry has reached the stage of “Talk is cheap, show me the product.” AI can write the code, but what truly has life is a product users want to use, are willing to pay for, and that can create value reliably.