Pioneers Insight Method Research Author
Vol. 52: A 74-Page PPT Breakdown of DeepSeek and AI Agents
Back to Episodes

Vol. 52: A 74-Page PPT Breakdown of DeepSeek and AI Agents

Summary

  • 庄明浩 believes the broad consensus in early 2025 is that leading vendors already sit between L2 “reasoners” and L3 Agents. The progression runs from ChatGPT-style conversation to o1, DeepSeek R1 and Kimi 1.5: models first had to “learn to think,” and now must “be able to act”; the “only conclusion today” is not that AGI has arrived, but that industry competition is shifting from model Q&A to real-world delivery. Model scores still matter for industry and investment debates, but application execution, workflows and organizational coordination are gaining weight.

  • DeepSeek’s breakout was not a single-point breakthrough by R1, but the result of technology, product, public sentiment and Lunar New Year emotions taking turns by date. V3 established the capability base on December 26, 2024; the App launched on January 10; R1 was released on January 20. The true product inflection came on January 24, when “Deep Think” and “Web Search” could first be enabled simultaneously. Only after that, alongside discussion in U.S. tech circles and 冯骥’s January 26 remark that it “might be a technology achievement tied to national destiny,” did traffic fully ignite.

  • The “DeepSeek was trained for $5.6M” comparison is seriously misleading, but the exponential decline in inference cost is real. The $5.6M covered only the compute cost of V3’s final successful training run, excluding hardware, personnel and prior investment; 庄明浩 cites U.S. estimates of tens of millions to $50M-$60M for a single training run at leading labs, implying a gap of several times to 10x rather than multiple orders of magnitude. Sam Altman’s rule of thumb is that usage costs fall 10x every 12 months; GPT-4 to GPT-4o fell 150x in roughly 18 months. Lower prices may expand total demand through Jevons’ paradox rather than compress it.

  • DeepSeek did not end compute investment; the 2025 capex narrative is accelerating. Microsoft, Amazon, Meta and Google are expected to lift related spending from roughly $222B in 2024 to about $320B in 2025; adding Stargate’s first-year $100B on 庄明浩’s estimate brings the total close to $420B. “More money—better infrastructure—better training—better models” is not stopping this year. Nvidia fell 17% in a single day on DeepSeek fears, wiping out roughly $560B, but recovered most of the loss in less than a month—evidence that the market has not abandoned the picks-and-shovels thesis.

  • The path that genuinely changed the technology narrative was the shift from human-labeled processes to outcome-rewarded reinforcement learning. Like AlphaZero, R1-Zero lets the model explore independently on objectively verifiable tasks, make mistakes and receive rewards or penalties only for the result. R1 then improved usability with a small amount of high-quality cold-start data, language-consistency rewards, general data and mixed rewards. If “all structured methods constrain model performance” becomes the dominant view, it would weaken the value of large-scale process labeling and put Scale AI-style business models under direct pressure.

  • Model access is rapidly becoming commoditized, forcing search, cloud and application companies to look for moats outside the model. 密塔AI搜索, 纳米搜索, 知乎直答, Perplexity, WeChat, 元宝, in-car systems and multidimensional spreadsheets can all connect to DeepSeek. Differences in information sources may not create durable barriers; two-button product designs and even the choice of whether to build an in-house model are converging. The most tradable metrics remain ARR and growth, while long-term defensibility may come from vertical industries, distribution, operations, brands or network effects—“money is still the bluntest standard,” just far from a complete answer.

  • Agent is not a category that has already converged, but a catch-all for whether a scenario, technology and product can close the loop. Once search or Coding delivers clearly, it gets called AI search, Cursor or Devin; capabilities whose boundaries remain unclear are still placed in the Agent bucket. A robot vacuum paired with a robotic arm shows the real difficulty: vision and reasoning models are only the starting point; classification, grasping, obstacle avoidance, pricing, incident handling and continuous iteration are what make a product. Until capabilities converge, “be cautious about building applications” and “the model is the application” will continue to pull in opposite directions.

Deep dive

1. The industry’s 2025 map now sits between L2 and L3

  • Roughly one-quarter of this 74-page PPT carries over the 2024 summary; the remaining three-quarters were rewritten in February 2025. DeepSeek changed the landscape so quickly that when 庄明浩 revisited the old draft, he felt much of it already belonged to “a previous era.”

  • Under OpenAI’s five-level roadmap, L1 is the chatbot, L2 is the “reasoner” that can solve problems, L3 is the Agent that can both think and act, L4 is the “innovator” that assists with invention and creation, and L5 can organize and complete entire bodies of work.

  • 庄明浩 places o1, DeepSeek R1, Kimi 1.5 and the reasoning models about to be released by various vendors at L2, with L3 as the next step: “The original basic conversational form is beginning to evolve into reasoning models, and then into true Agents built on those reasoning models.”

2. DeepSeek’s breakout was a date-specific chain of events

  • DeepSeek V3 launched on December 26, 2024, and was already highly rated in Chinese and U.S. technology circles; some media even called it “Pinduoduo for large language models.” The official App went live on January 10, but had not yet become a mass-market hit.

  • R1 launched on January 20, the same evening 梁文锋 attended an enterprise symposium chaired by 李强. 庄明浩 stresses that the invitation could not have been arranged at the last minute because of R1’s release that day; if anything, it shows that V3 was the first model to win serious recognition.

  • On January 21, Trump announced Stargate, a $500B AI infrastructure plan. On January 24, Marc Andreessen, Scale AI’s Alexandr Wang and other U.S. entrepreneurs and investors began discussing DeepSeek intensively, after which domestic media brought the story “back home” through re-exported coverage.

  • At 23:32 on January 26, 冯骥 publicly recommended DeepSeek. On January 27—庄明浩 added “if I remember correctly”—Nvidia fell 17% in a single day, erasing about $560B in market value. On January 28, Lunar New Year’s Eve, a fabricated internal letter from 黄仁勋 and 梁文锋’s Zhihu reply completed the second wave of distribution.

3. V3, R1 and the full product experience created the differentiation

  • Kimi, MiniMax, StepFun and Qwen all released models around the same period, and Qwen was also open source with strong leaderboard results. 庄明浩’s central question is: “Why was DeepSeek the only one to go viral?” The answer cannot be reduced to either model strength or open source alone.

  • V3 and R1 ranked high enough to approach o1 and even o3 on some benchmarks. At the same time, DeepSeek was packaged as a Chinese model, an open-source model and an “extremely cheap” model, allowing technical capability to arrive alongside national narrative and cost shock.

  • Before January 24, “Deep Think” and “Web Search” could not be enabled at the same time: the former displayed a chain of thought but was limited to knowledge through December 2023, while the latter connected to the web but used V3 to answer quickly. Once both buttons could be pressed together, reasoning, real-time information and visible process became a complete product experience.

  • This also explains why the App launched on January 10 without immediately exploding: the broader public was seeing the reasoning process for free for the first time, while allowing the model to keep thinking against the latest information—not merely switching to another chat interface.

4. The $5.6M was the cost of one successful training run, not total company investment

  • The roughly $5.6M disclosed in the V3 paper comes from approximately 2.8M training hours multiplied by the unit price, and refers to the compute cost of the final successful run. 庄明浩 cautions that it excludes hardware purchases, personnel and prior R&D investment, so it cannot be directly compared with OpenAI’s annual losses or Stargate’s total investment.

  • Based on media estimates he cites, the single-run training cost for leading models at OpenAI, Anthropic and Meta is broadly in the tens of millions to $50M-$60M. The gap may be several times, or even 10x, but not the multiple orders of magnitude implied by public hype.

  • The decline in usage costs is more certain. Sam Altman says the cost of calling a comparable frontier model falls roughly 10x every 12 months; GPT-4 to GPT-4o fell 150x in about 18 months. “Two years is 100x. Two years is not 20x.”

5. The “national destiny” narrative amplified both traffic and the boundary between fact and fiction

  • 冯骥’s recommendation hit five points at once: strong capability, low cost, a permissive open-source license, free access and web connectivity, capped by the claim that “DeepSeek might be a technology achievement tied to national destiny.” From that point, it was no longer merely a model company, but a representative player in the U.S.-China AI competition.

  • Two AI-generated fake articles captured the paradox of this distribution cycle. 庄明浩 concluded that a Zhihu reply was fake and contacted the platform to remove it, yet people still responded: “But it was written really well. The emotion was exactly right.” Once the screenshot began circulating, it had detached from the original account.

  • On the first day of the Lunar New Year, he asked: “If this develops for several more years, who will be real and who will be fake? If the vast majority of people believe the fake thing, is it still fake?” DeepSeek’s breakout and the flood of AI-generated misinformation became two opposing faces of capability demonstration on the same day.

  • As of February 24, he estimated that DeepSeek’s web traffic and App DAU may have fallen roughly 70% from their peaks. Some of that reflects the team’s failure to actively capture the traffic, but “how much of this windfall is left” has become a more realistic question than download rankings.

6. DeepSeek’s hardest questions have shifted from training to strategic trade-offs

  • The first trade-off is safety. There are real shortcomings in model alignment, App and web data handling, and adaptation to AI policies in different countries. 庄明浩 rejects the idea that every criticism is an enemy attack, but there is no simple answer to how far the response should go.

  • The second is To C: do the App, Web and this wave of individual users actually matter? If DeepSeek captures the traffic, it must decide whether to charge, operate the user base and enter vertical scenarios; if it does not, it must explain what product form the model is meant to take in the future.

  • The third is To B. The broad MIT license reads like “use it if you want,” and DeepSeek even closed its API top-up channel at the peak. But not prioritizing monetization does not mean commercial pricing can be deferred forever.

  • Financing is more difficult still: what would a purely financial investor be for, and the value of dollar-denominated VC is limited. Taking strategic investment from Alibaba, Tencent or Huawei would amount to choosing sides. Whether to accept “national team” capital such as the National Social Security Fund is, in 庄明浩’s words, a “very, very difficult question,” so he is choosing to watch and wait.

7. The dominant narrative of the past 2 years was still capital-driven scaling

  • 庄明浩 compresses the industry logic from late 2022 through late 2024 into 4 steps: “More money, better infrastructure, better training, better models.” OpenAI’s valuation rising to $155B is the clearest expression of that progression.

  • The combined market capitalization of the 7 U.S. tech giants rose from $11.8T at the start of 2024 to $17.6T by year-end, peaking at roughly $18T. They contributed more than 55% of the S&P 500’s full-year gain; all other companies together contributed about 45%.

  • Nvidia rose from roughly $1.2T in 2023 to more than $3T, benefiting from 3 waves of upside in gaming, crypto mining and AI. After plunging 17% in a single day amid the DeepSeek shock, it recovered most of the loss in less than a month. Data-center operations may already account for roughly 70%-80% of revenue.

8. Cloud, capex, chips and energy are still expanding in sync

  • AWS, Microsoft Azure and Google Cloud could still maintain quarterly growth of 35%-40% around 2020; by 2022-2023, growth had fallen to the teens. After ChatGPT launched, growth returned to above 20%-25% from the second half of 2023.

  • Capex at Microsoft, Amazon, Meta and Google is expected to rise from roughly $222B in 2024 to about $320B in 2025, a 44% increase. The individual expectations are approximately $80B, $100B, $60B and $70B-$75B, respectively.

  • Adding Stargate’s first-year $100B on 庄明浩’s estimate would bring related investment close to $420B, nearly double the 2024 level. “In the adult world, money is the simplest standard”: tech giants operate in the tens and hundreds of billions, major players in the tens of billions, and challengers in the hundreds of millions.

  • The giants are buying Nvidia chips while developing their own to avoid being “held by the throat” by a single supplier. Data centers are also hitting an energy ceiling, making nuclear power the emerging consensus direction. 庄明浩 cites forecasts that by 2030, U.S. data-center electricity consumption could exceed the full-year consumption of Japan or the UK.

9. Ecosystem investment and China’s “copying the homework” follow the same capital path

  • U.S. giants are pulling challengers into their ecosystems through investment and acquisitions. 2024 saw “hollowing-out acquisitions”: Character.AI was sold to Google, the Inflection team went to Microsoft, the Adept team went to Amazon, and 01.AI handed its team to Alibaba.

  • The giants’ own businesses are also part of the picture: Nvidia chips, Microsoft Copilot, Apple Intelligence, Google models, Amazon Nova, Meta Llama, Tesla FSD and xAI all serve the closed loop of “infrastructure plus proprietary applications.”

  • Alibaba, Baidu, ByteDance, Huawei and Tencent in China likewise cover models, cloud, applications and multimodality. Alibaba says its infrastructure investment through Alibaba Cloud over the next 3 years will exceed the previous 10 years combined; 庄明浩 believes that could mean more than RMB100B per year, using an estimate of roughly RMB120B-RMB130B as an example—barely reaching the spending threshold of U.S. tech giants.

  • Startups have already split into different paths: 01.AI is shifting toward To B implementation, Baichuan toward healthcare, and StepFun and Zhipu are respectively tied to Shanghai and Beijing. MiniMax and Moonshot AI continue to pursue frontier foundation models, but DeepSeek has sharply increased the strategic pressure on that route.

10. Jevons’ paradox offers the counterargument to whether low prices destroy compute demand

  • On the day Nvidia collapsed, visits to Wikipedia’s “Jevons’ paradox” page rose roughly 200x. The historical analogy—improvements in coal efficiency ultimately expanded coal use—was applied to the question of whether falling model costs would hurt suppliers or unlock more AI demand.

  • Pretraining data could approach its limit around 2028 in an extreme case; humanity may simply have no more data to feed large models. Ilya Sutskever once said, “If you train a huge neural network on a huge dataset, it will inevitably succeed.” 庄明浩 then argued that the formula has now failed: pretraining Scaling Law can no longer explain progress by itself.

  • In June 2024, Claude 3.5 improved sharply in copywriting and Coding, and Anthropic internally also viewed it as a reasoning model. OpenAI released o1 in September, as reinforcement learning, post-training, math and programming began to form a new technical paradigm together.

11. R1-Zero showed that outcome rewards do not require process imitation

  • The early route to reproducing o1 was to make the model “slow down”: use longer chains of thought to stimulate more internal information, then reward each reasoning step through a PRM. If that route worked, demand for labeled data would grow with every Step, naturally benefiting Scale AI.

  • The conclusion from a Kimi employee’s public retrospective was the opposite: “With precise rewards, do not use structured methods. Ultimately, all structured methods will constrain the model’s performance.” The model should search independently, be allowed to make mistakes and receive rewards only for the final verifiable result.

  • 庄明浩 uses AlphaZero to explain R1-Zero: give the model only the rules of Go, let it play against itself, and reward wins and penalize losses; it can escape human game records and preferences. Applied to reasoning training, this means letting V3 perform pure reinforcement learning on objective, measurable tasks—“truly breaking free from human constraints.”

  • R1-Zero still mixed Chinese and English, produced excessively complex chains of thought and performed poorly on general Q&A. R1 therefore added a small amount of high-quality cold-start and chain-of-thought data, supervised fine-tuning, language-consistency rewards, and a mix of reasoning and general data for final fine-tuning and reinforcement learning. 庄明浩 views the early language mixing in o3-mini as a similar phenomenon.

12. Distillation means DeepSeek exported not just a model, but a method

  • DeepSeek applied the same training method to distill Qwen and Llama models of different sizes, and paper evaluations showed that these models also improved. The key point is that the methodology is usable not only by DeepSeek, but by other open-source models as well.

  • The many “locally deploy R1” tutorials online are in practice deploying these smaller distilled models, not the full-strength R1.

  • This also explains why Alexandr Wang forcefully elevated DeepSeek into a U.S.-China competition issue: if outcome rewards and self-exploration reduce the need for large-scale human labeling, Scale AI’s addressable market will shrink. 庄明浩’s point is not that every safety concern is wrong, but that commercial interests cannot be separated from the public positioning.

13. AI search is becoming a competition between 2 buttons and the same pool of models

  • DeepSeek now has 4 modes: pure conversation, web search, deep thinking, and both buttons enabled together. The last is the most important because it combines real-time information, reasoning and visible chains of thought; previously, o1 required paid access, and the $20 price threshold was still high for many users.

  • 密塔AI搜索, 360纳米搜索, 知乎直答 and Perplexity have all connected to DeepSeek. Whether differences in information sources can create a sufficient moat is now doubtful, while product design has largely converged on the same 2 buttons. At this point, there is no longer an obvious answer to whether an AI search company should build its own model.

  • Anker’s product page for a 140W charger directly displays answers from DeepSeek and 豆包 under the headline “The charger brand most recommended by leading AI,” with the labels “No data modified,” “Answered directly without training,” and “Actually tested.” A ramen shop in Shanghai also turned a DeepSeek recommendation into a sign at its entrance.

  • 元宝’s answers include links to 58.com, 易家宝 and app downloads, prompting users to think of SEO again. 庄明浩 says no one is specifically doing this yet, but once brands start treating AI answers as a standard, “the results generated naturally today” could evolve into a new optimization industry aimed at AI search.

14. DeepSeek’s spillover traffic reconnected cloud, devices and distribution

  • DeepSeek could not absorb peak demand, making cloud and inference providers the direct beneficiaries. SiliconFlow saw traffic surge and raised financing; a snapshot cited by 庄明浩 even showed its traffic temporarily exceeding Tencent Cloud’s—“simply unimaginable for a startup.”

  • The integration range expanded from Xueqiu to Huawei Xiaoyi, OPPO Find N5, BYD’s in-car system, Feishu Bitable and 文心一言, and then to WeChat. Xinhua specifically confirmed that WeChat was testing the integration, prompting 庄明浩 to joke: “A national-destiny-scale model plus a national-scale WeChat application naturally requires the national news agency to report it.”

  • One path for Douyin creators is to use Coze to generate AI avatars, then distribute them across search, livestreams, group chats and comment sections for Q&A or selling products. This embeds model capability directly into existing content and transaction networks.

  • When every company can connect to the same model, the moat has to be reconsidered: product execution, user operations, brand, network effects or vertical industries? The U.S. private market still looks first at ARR and growth. 庄明浩 calls Cursor “the AI company that should reach $100M ARR the fastest,” adding that “money is still the bluntest standard.”

15. The real Agent threshold is product closure, not a new label

  • Coding shows the progression of capability: ChatGPT can generate code from prompts but cannot debug end to end; GitHub Copilot understands code and offers suggestions; Cursor adds logical understanding, file operations and command-line execution; Devin then triggered debate over whether junior programmers are still necessary.

  • Search systems break down tasks, call information sources, organize results and present answers, so they can be called Search Agents. Coding systems understand a goal and execute continuously, so they can be called Coding Agents. 庄明浩’s sharper definition is that once the scenario and delivery are clear, it is called AI search or AI Coding; what remains unclear is grouped under Agent.

  • A traditional Agent needs at least planning, memory, tool calling and communication with other Agents. Since 2024, construction has shifted from code and APIs toward window-based drag-and-drop and workflows; tasks have moved from single-threaded execution to Multi-Agent systems, while evaluation has shifted from the model itself to “the ability to solve problems.”

  • A robot vacuum paired with a robotic arm makes the abstract problem concrete: a mature vision model is only the starting point. The system still has to distinguish trash, shoes and dirty laundry, handle weight, grasp objects, avoid obstacles, provide proactive service, set pricing, manage accidents and iterate. The difficulty is the complete chain of “defined scenario—technical boundary—product delivery—updates and iteration.”

  • Organizations are beginning to split along the same chain. Alibaba moved the Tongyi App into an intelligent information business group that includes Quark and UC, while leaving the model in Alibaba Cloud. Baidu and ByteDance have separated application and model teams; Tencent placed 元宝, Sogou, Sogou Input, QQ Browser and ima under CSIG, while model research remains in TEG.

  • But “the model is the application” continues to challenge this division of labor. One view is that DeepSeek introduced almost no product-form innovation, and that its breakout was driven primarily by intelligence and open source. AI-native social products requiring real-time text, voice, video and lip-syncing are still judged to be 2-3 years away. 李翔’s closing reminder follows: “Until capabilities converge, be cautious about building applications.”