Vol.51 The Rumors and Misconceptions About DeepSeek
Summary
- The global breakout of DeepSeek R1 was not a victory on a single benchmark, but the result of a technology breakthrough, the R1-plus-web-search product combination, and US-China narratives resonating within days. 张涛 observed that roughly “80%” of the impressive cases posted by overseas users had search enabled, giving many people their first experience with a model that searches, reasons and reflects in real time. On January 24, Marc Andreessen moved from reluctant acknowledgment to calling it “the best gift humanity has received”; on January 26, 冯骥 framed it in terms of “national destiny.” NVIDIA then plunged 17%, turning a technical event into a mass public and capital-markets event.
- V3 solved how to train a first-rate base model with constrained compute, while R1 publicly demonstrated a reinforcement-learning path driven by outcomes. V3 used 2,048 H800s to produce what 张涛 called a GPT-4o- and Claude 3.5-level model, relying on a large number of “ingenious hacks combining engineering and algorithms”; R1 then used the ORM route to challenge the PRM route that had dominated the industry’s attempts to replicate o1. 张涛’s standard was blunt: “If you don’t publish a paper and don’t release the model weights, then as far as we’re concerned, you don’t have it.”
- “R1 was trained for $6M” is a myth manufactured by the communications chain; the original figure was merely the auditable GPU cost of V3’s final training run. The technical report cited roughly 2.8M H800 GPU hours, which at $2 per hour works out to more than $5M, or roughly $6M, while explicitly excluding prior research, experiments, architecture exploration, data preparation and cleaning. The fair comparison is a roughly $6M single training run versus tens of millions for other models—not that figure versus Meta’s total investment or Anthropic’s funding or valuation. 张涛 believes “DeepSeek never intended to deceive anyone.”
- R1 has challenged NVIDIA’s scarcity narrative in the short term, but could expand inference demand by two orders of magnitude over the long term; the incremental demand may simply no longer belong to NVIDIA alone. 张涛 expects inference compute demand to grow 100x within 3 years, summarizing the industry-versus-Wall-Street divide as “everyone on the West Coast is buying, everyone on the East Coast is selling.” 庄明浩, meanwhile, noted that NVIDIA’s gains and losses come from the same source: it is both the biggest beneficiary of the AI wave and the most volatile name when sentiment reverses. Cases of R1 running on Huawei 910C and 昇腾 suggest that the market is repricing not total compute demand, but the share structure of inference chips.
- Some argue that US chip controls forced Chinese engineers to innovate, but claims that DeepSeek secretly used 50,000 H100s and that its model was distilled based on identity responses both lack reliable evidence. The paper explicitly names H800s, while 张涛 drew a straightforward conclusion from the interconnect-bandwidth hacks in V2 and V3: “They really didn’t have the chips.” Dylan Patel initially said 50,000 Hopper GPUs, which Alexandr Wang later relayed as 50,000 H100s. Distillation is widespread across the industry, but the model occasionally identifying itself as ChatGPT is more likely a result of pre-training data and insufficient self-cognition alignment; it does not prove shell use, infringement or distillation.
- 张鹏 believes DeepSeek’s biggest shock to the US is that it showed the existing AI policy of “accelerating America and slowing China” failed to create the expected gap, making subsequent restrictions potentially more investment-relevant than the technical debate itself. Policy tools already cover GPUs, semiconductor equipment, sensitive data, outbound investment and potentially talent; the Trump administration’s first-100-days design, whether H20 faces further controls, and Josh Hawley’s proposed US-China AI decoupling bill all merit watching. Europe is more likely to deal with privacy and compliance case by case, while the US may invoke national security as it did with TikTok: “No amount of compliance work may be enough.”
- For Chinese AI companies, R1 shows that chatbot traffic is not a moat; the opportunity is shifting toward RAG, tool use and Agents built on open-source models. 豆包 reached roughly 20M DAU after more than a year of promotion, while Kimi never reached 10M; DeepSeek exceeded that scale within days, yet may not even treat DAU as a KPI. 庄明浩 said “an anomaly has appeared at the table,” forcing Kimi, MiniMax and others to revisit open source, closed source and commercialization. For enterprises, 张涛 advises against further pre-training: use existing models, RAG, targeted post-training and private deployment instead. 庄明浩 relayed one company leader’s view that traditional businesses cannot simply wait for the technology to stabilize; they need to engage deeply now.
Deep dive
1. DeepSeek Fermented in Tech Circles Before US-China Narratives Turned It Into a Mass Event
Lily’s test for whether DeepSeek had “broken out” came from two non-technical signals: first, 和菜头 wrote a dedicated post expressing his shock; then her nearly 40-year-old sister, who “doesn’t understand AI at all,” proactively asked about DeepSeek. To her, that was stronger evidence than any exaggerated public-account headline that the story had crossed beyond the industry.
张鹏 described the distribution path as “a bit like exporting and then re-importing”: after V3 and R1 launched, mainstream US media showed little reaction for the first few days, partly because Trump’s inauguration dominated the headlines and partly because the US side was still working out what the models meant for the American AI industry and US-China competition. Once The New York Times published, coverage became blanket.
庄明浩 recalled that V2 had already generated substantial discussion in open-source and technical circles, while V3 spread into technology, product and internet communities. During a series of livestreams after Trump’s inauguration, he discussed the impact of US-China technology competition and “Stargate,” mentioning R1’s challenge to the high-capex AI narrative along the way—but he did not expect the story to turn into a nationwide celebration.
张涛 saw demand emerging even earlier at the product level: around January 23, sophisticated users in China and overseas were already asking whether the products they used had integrated R1, and explaining through their own workflow examples why it was different. Real usage drove the industry’s initial reaction; the broader public breakthrough came later.
2. “National Destiny,” the Market Crash and Influencer Sentiment Delivered the Final Spark
张涛 identified the US turning point as Marc Andreessen’s entry on January 24. He initially said the model was genuinely impressive but “I’m not happy that it’s impressive,” then quickly shifted to calling R1 “the best gift humanity has received” and describing it as AI’s “Sputnik moment.”
Andreessen had previously taken a distinctly confrontational stance toward China, making his endorsement more transmissible than that of an ordinary technology KOL. 张涛 admitted he was “a little stirred up,” because it meant work by a Chinese team had been recognized by an American opinion leader who was not politically sympathetic.
庄明浩 placed the domestic turning point on the evening of January 26, when 冯骥 posted about DeepSeek and elevated it for the first time to the level of “national destiny.” NVIDIA then plunged 17%, translating a technical issue that still required explanation into an extreme market move visible to everyone.
张涛 is also a devoted user of Black Myth: Wukong, so he described the resonance between 冯骥 and DeepSeek as “a double-fandom celebration.” From that point on, product capability, national sentiment, geopolitics and US equity positioning could no longer be cleanly separated.
3. AI-Generated Fake Letters and Responses Showed Consensus Moving Faster Than Facts
After NVIDIA’s plunge, a text claiming to be an internal letter from 黄仁勋 went viral in private social feeds. On January 28, a Zhihu article supposedly written as 梁文锋’s response to 冯骥 was widely reposted. 庄明浩 confirmed that both were AI-generated and personally reported the account impersonating 梁文锋 to Zhihu’s operators.
Lily had also praised the supposed “梁文锋 response” in a group chat for being moving, only to learn from 庄明浩’s social feed that it was fake. She saw no obvious AI fingerprints in the language; the mistake gave her a visceral first-hand sense that “this really is different from before.”
After repeatedly warning friends, 庄明浩 wrote a sharper question: “When a consensus forms, and even if it is false, many people believe it, is it still false?” The reasoning advances represented by R1 have arrived just as humans are finding it harder to determine who wrote a moving passage.
张鹏 added that policy circles had their own fakes, including a supposed David Sacks article outlining a “five-step knockout” of China’s AI industry and sensational claims that anyone downloading DeepSeek would be fined $100M. Both spread by exploiting policy complexity and collective anxiety; readers often had never seen the original interview or bill.
4. V3’s Breakthrough Was Training a First-Rate Base Model on Constrained Hardware
张涛 summarized V3’s core achievement as a large number of “ingenious hacks combining engineering and algorithms,” stressing that he meant it positively: with limited compute, the team used 2,048 H800s to train what he described as a GPT-4o- and Claude 3.5-level base model.
The innovation came from combining engineering with algorithms and cannot be reduced to the label “cheap.” Model size, active parameter count, training data and GPU hours can be reconciled against one another, so what the industry first saw was a verifiable systems-engineering achievement.
But if one looks only at model capability, 张涛 believes V3 may still feel to ordinary users like “another GPT-4o, another Claude.” Its significance inside the industry was enormous, but it could not by itself explain how a model entered family New Year conversations and global political debate.
5. R1 Publicly Proved the ORM Route and Changed the Methodology of Reasoning Models
张涛’s technical assessment is that the industry’s earlier attempts to replicate o1 followed the PRM route, rewarding the reasoning process step by step. R1 was the first public demonstration that a more direct ORM route—reinforcing the final outcome—could work.
The idea was not unprecedented; what had been missing was a public demonstration of how to make it work. 张涛 acknowledged that OpenAI may have achieved it earlier, but the open-source standard was clear: “If you don’t publish a paper and don’t release your model weights, then as far as we’re concerned, you don’t have it.”
R1 therefore mattered as more than a new model. It gave the industry a working example: V3 supplied the base model, reinforcement learning produced visible reasoning capability, and the paper and weights allowed others to continue experimenting along the same paradigm.
That is also why R1 cannot be conflated with its smaller distilled variants. The original R1 depends on both V3 and reinforcement learning; a small model trained only on SFT data generated by R1 has not inherited the full training path.
6. What Made the Public Say “Wow” Was R1 and Web Search Working Together
Before R1, only a minority of users had actually used a reasoning model. o1 required ChatGPT Plus or Pro, which 张涛 cited as $20 and $200 per month respectively, and at the time o1 could not call search for real-time information.
After reviewing impressive cases posted by users worldwide, 张涛 found that “80%” had search enabled. Users assumed the achievement came from the R1 model alone, but the actual experience was a loop in which the model retrieved real-world information, continued reasoning and then reflected and corrected itself. He repeatedly emphasized: “This is a product, not a model.”
庄明浩 filled in the product-evolution detail: the DeepSeek App launched in mid-January, but initially the “R1” and “web search” buttons could not be selected at the same time. Without search, R1’s data roughly stopped at December 2023. Seven or eight days later, the two buttons could be enabled together, and the experience changed qualitatively.
Lily’s usage feedback confirmed the point. When the servers were smooth and web search worked, answer quality exceeded anything she had experienced with OpenAI Plus or Claude; once web search became unavailable, answers to similar questions deteriorated sharply. 张涛 therefore believes that releasing R1 without search might not have produced anything close to the same breakout.
7. The $6M Was Only the GPU Bill for One V3 Training Run
The “$6M” was not the full cost of R1; the figure came from V3’s technical report. The report cited roughly 2.788M H800 GPU hours, which at a rental cost of about $2 per hour works out to more than $5M, or roughly $6M.
张涛 emphasized that the report provided more than a total: it broke out the costs of pre-training, context expansion and post-training, and disclosed model size, active parameters and data volume. The industry can back into training time from the formulas, so there is little dispute over the GPU bill itself.
The same table explicitly states that this covered only V3’s final training run, excluding earlier research, failed experiments, algorithm and architecture exploration, data preparation and cleaning. Adding those exclusions back in gets closer to DeepSeek’s or 幻方’s actual long-term investment.
The formulation accepted by both 张涛 and 庄明浩 was therefore: V3’s single-run training cost was materially below that of comparable models, but “more than $5M to roughly $6M” absolutely does not mean the company spent only that amount to build R1 from scratch.
8. Media Inflated a Single Training Cost Into a Commercial Myth
The original technical comparison was $6M versus tens of millions for a single Llama training run. 张涛 considers that entirely reasonable and sufficient to show that DeepSeek saved substantial costs. The distortion began when non-technical KOLs and traditional media entered the story.
The first distortion compared $6M with Meta’s entire investment in Llama, putting multiple rounds of experiments and total spending on the other side. Others then compared it with companies such as Anthropic’s funding, and even their valuations. At that point, the two sides were no longer measuring the same kind of cost.
Once the myth took hold, the counter-narrative branded DeepSeek a fraud. 张涛’s target was not the technical report but the communications chain: “DeepSeek never intended to deceive anyone.” Most exaggerated versions were manufactured by English-language media and KOLs lacking basic industry knowledge.
庄明浩 explained the story’s reach with one line: “In the adult world, money is the simplest yardstick.” When mainstream AI news is usually priced in billions, tens of billions, hundreds of billions or even $1T, the sudden appearance of “$6M” hits ordinary people through the difference in zeroes before they need to understand the architecture.
9. The Gap in Capex Scale Magnified the DeepSeek Shock
庄明浩 divided industry participants into 3 tiers: technology giants spending tens of billions of dollars annually on AI infrastructure; major players such as OpenAI, Anthropic and xAI spending in the billions; and challengers spending hundreds of millions. Against that scale, $6M looks almost like a statistical error.
Recalling financial statements on the program, he cited roughly $80B in related spending for Microsoft the following year, $65B for Meta, $75B for Amazon and $70B for Google. The top 6 technology giants spent roughly $230B combined the prior year, a figure that could rise to $300B-$400B.
The “Stargate” plan created a grand narrative around $500B in total spending; adding the $100B discussed within it would push the figure even higher. Even if a comparable model typically costs only $20M-$30M to train once, the news environment still portrayed DeepSeek as using loose change to overturn the entire capex system.
张鹏 noted that people in both countries had incentives to amplify the story: in the US, to question whether OpenAI’s policy support, capital access and export controls were effective; in China, to produce a breakthrough narrative about overtaking the US at low cost. When emotion is being amplified on both sides, “the facts become less important.”
10. The English-Language Information Gap Created 3 Categories of Company Rumors
张鹏 believes US media had long ignored Chinese AI, leaving English-language information “extremely behind and delayed” when DeepSeek suddenly needed coverage. That produced not only cost misunderstandings, but also persistent errors about company facts that are easy to verify in Chinese.
The first category cast the young researcher 罗福莉 as DeepSeek’s “secret weapon,” ignoring Chinese reports that she had already moved to Xiaomi and other details. The second portrayed DeepSeek as a “side project” accidentally created by a quantitative hedge fund.
To Chinese AI professionals, DeepSeek was plainly not a side project undertaken on a whim. It was a company being built deliberately, yet the narrative of a quantitative fund accidentally creating the world’s best AI continued to circulate in English-language media.
The third category rewrote V3’s final-run $6M as the entire R1 R&D budget. Together, the 3 stories created a highly transmissible legend while obscuring the organizational capabilities and engineering path that actually deserved study.
11. DeepSeek’s Outage Accidentally Educated the Entire Open-Source Reasoning Ecosystem
DeepSeek lacked the operational capacity to handle the flood of consumer traffic, so users began searching for local or third-party cloud deployments. 庄明浩 believes this created previously unimaginable growth for domestic inference infrastructure providers, chatbot companies and model-hosting platforms.
硅基流动 did not immediately commit all its resources when V3 and R1 first launched; it accelerated only after the breakout. Large numbers of users then registered, requested APIs, and learned how to call and deploy models. Invitation rewards of about RMB14 quickly filled Xiaohongshu comments and KOL posts with referral codes.
Bilibili’s homepage simultaneously featured at least 3 tutorials for local or cloud deployment, with 庄明浩 saying each video had roughly 10,000 viewers online at once. APIs, deployment and the surrounding division of labor—previously known only to experienced professionals—entered the field of view of ordinary users.
The long-term value of this ecosystem education was impossible to estimate at the time. At a minimum, the wave exposed more people to the relevant vendors, services and deployment methods.
12. NVIDIA’s 17% Drop Was a Crowded-Narrative Detonation, Not the Pricing of One Fact
庄明浩 believes NVIDIA’s 17% drop cannot be attributed to a low-cost model alone. Stocks are the result of competing positions; existing shorts, valuation disputes and accumulated sentiment all needed a trigger, and DeepSeek ignited them simultaneously.
He described NVIDIA as a case of gains and losses coming from the same source: it was the biggest beneficiary of the foundation-model wave over the previous 2 years and is also one of the Magnificent 7, concentrating the most positioning and debate. He had previously compiled the S&P 500’s biggest one-day movers and found that roughly 16 or 17 of the top and bottom 20 names were AI-related.
The market then formed 2 opposing logics: low costs weaken chip scarcity, or low costs lower the application barrier and expand total demand. Neither side could persuade the other, so the views continued to hedge one another through the prices of NVIDIA, cloud providers and application companies.
The same event produced differentiated pricing. Apple did not fall when NVIDIA plunged; edge-computing logic then drew attention to Lenovo and Xiaomi; Google still fell roughly 8% after earnings despite presenting aggressive AI capex. The same information does not mean the same thing for every company.
13. Compute Demand Could Grow 100x Over the Long Term, but the Incremental Demand Need Not All Go to NVIDIA
张涛 summarized the market divide through an observation relayed by 真格基金’s 雨生: “Everyone on the West Coast is buying NVIDIA; everyone on the East Coast is selling NVIDIA.” The industrial side believes open-source breakthroughs will expand training and inference, while financial markets first trade the possibility of lower returns on capex.
As an industry participant, 张涛 is explicitly bullish on total demand. He believes inference compute could be 100x larger within 3 years—“within 3 years,” not 3 to 5 years—and not merely 10x. Cheaper models do not necessarily mean less total computation.
What truly changes is the share structure. 硅基流动 has already deployed R1 on Huawei 910C and the 昇腾 ecosystem, naturally prompting the market to ask whether other chips can also handle inference. Before R1, the lack of an open-source model able to compete head-on with leading closed models weakened the incentive to deploy alternatives.
R1 created large-scale, portable real-world deployment demand for the first time, and showed users that “maybe you don’t need NVIDIA after all.” Long-term compute demand can therefore rise while NVIDIA’s share of that demand is no longer guaranteed to remain exclusive.
14. “50,000 H100s” Came From Turning the Hopper Category Into a Specific Model
张鹏 traced the claim back to Dylan Patel, who said in November 2024 that DeepSeek had more than 50,000 Hopper GPUs—not that they were all H100s. H800s are also Hopper GPUs, but with memory bandwidth cut back under US restrictions.
During Davos, Scale AI CEO Alexandr Wang directly changed the claim in an interview to 50,000 H100s, inferring that the chips violated export controls and that the company therefore dared not disclose them. The claim prompted attention and investigative mechanisms from the US Commerce Department, the White House National Security Council and others.
张鹏’s interim judgment is that 50,000 H100s is probably false and that DeepSeek relied primarily on H800s; the V3 paper also explicitly names H800s. The program offered no definitive conclusion on whether the investigation would continue or whether the chips came from other sources.
Policy interpretations then split. One camp said export controls had forced Chinese engineers to optimize their systems to the limit, making the US “suffer the consequences of its own policy.” The other said that if H800s could train the models, the cut-down chips should also be banned, making H20 a possible target for the next round of controls.
15. Restricted Optimizations in V2 and V3 Were Used to Support the Compute-Constrained View
The technical analysis cited by 张鹏 said the team addressed H800’s limited memory bandwidth by dedicating roughly 20 of each chip’s 132 processing units to managing cross-chip communication, while using the lower-level PTX instruction set rather than relying only on the standard CUDA path.
According to the analysis, such optimizations are valuable only when communication and bandwidth are constrained. While reading the V2 and V3 papers, 张涛 kept asking, “Why would they do this?” If the team really had the large supply of full-strength H100s claimed by Alexandr Wang, it would not have needed so many hack-like solutions.
庄明浩 explained 幻方’s early access to compute. Quantitative trading already depends on machine inference, computation, analysis and summarization; before the US restrictions fully took effect, 幻方 had accumulated a substantial number of cards while developing both quantitative strategies and foundation models. That does not prove R1 used unauthorized H100s.
张涛 also cited an estimate possibly from SemiAnalysis: roughly 60,000 cards, including 10,000 A100s, 10,000 H100s, 10,000 H800s and 30,000 H20s. He considered it a relatively credible external estimate, but found the engineering trade-offs in the papers more intuitive: “They really didn’t have the chips.”
16. Distillation Is a Common Training Method, Not “Stealing the Essence” of Another Model
张涛 began with the stricter historical definition: a large teacher model generates outputs, then a smaller student model with fewer layers or parameters from the same network family is trained on those outputs to approximate the teacher’s performance.
The definition later broadened: teacher and student can use different architectures, as long as the large model’s outputs guide the smaller one. As 张涛 summarized it, distillation generally runs from large to small, and in theory the student should not surpass the teacher through distillation alone.
The technique is widespread in both the Chinese and US industries. “Behind closed doors, everyone knows it; in public, nobody admits it.” Even if DeepSeek used outputs from other models, their role in V3 and R1’s full training chain could only have been a local component, not a substitute for architecture, data and reinforcement learning.
Imagining distillation as boiling soup until “the water is gone and only the essence remains” naturally creates a sense of plagiarism. The real process is training another model on output data, with what it can learn constrained jointly by the teacher’s outputs, data design and the student’s capabilities.
17. R1 Calling Itself ChatGPT Does Not Prove Shell Use or Distillation
The most common evidence for the distillation accusation is that DeepSeek occasionally answers “I am ChatGPT,” or says in reasoning tokens, “As OpenAI’s ChatGPT, I cannot…” 张涛 believes that inferring shell use or theft from such screenshots reveals a lack of understanding of how modern LLMs are trained.
During pre-training, a language model does not know who it is. Self-cognition is part of post-training alignment and requires dedicated instruction about the model’s identity. Internet corpora contain large numbers of lines saying “I am ChatGPT,” and even slightly incomplete data cleaning can make a model reproduce similar language.
张涛 believes R1 paid far less of the “Alignment Tax” that closed commercial models must bear. Safety, harmlessness and identity alignment can make a model “dumber” along certain dimensions; the R1 paper even notes that on one benchmark, omitting the harmlessness data used for alignment could improve results by roughly 7 to 8 points.
If someone really wanted to distill a model, the most valuable asset would be the full probability distribution from the teacher’s next-token prediction each time—not merely the final token. The OpenAI API had already blocked those detailed probabilities more than a year earlier. Trainers would also not deliberately instruct every answer to begin with “I am ChatGPT,” so identity errors are not valid evidence.
18. Model Theft Has Clear Boundaries, While Distillation Infringement Remains a Legal Gray Area
The clear cases of “stealing a model” listed by 张鹏 include physically taking storage devices, breaking interface or endpoint security protections, and hacking into networks to obtain unreleased weights. Beyond civil liability, these actions may trigger administrative or criminal liability under cybersecurity laws.
Another category is shell use: legally obtaining a model, introducing no new data and making no material changes to its architecture, training, fine-tuning, alignment or inference, while falsely claiming it was independently developed. That can involve copyright, fraud and open-source license violations, making the legal analysis relatively clear.
If a model is independently developed but obtains key data mixes, architecture parameters or performance optimizations through unknown channels, it may infringe trade secrets. Knowledge distillation itself is not automatically illegal and does not automatically infringe intellectual property; where an agreement expressly prohibits it, the first issue may instead be breach of contract.
张鹏 therefore rejects the jump from a few output screenshots to “stealing US intellectual property.” Model distillation will become a new AI intellectual-property research issue, but it requires rigorous proof of facts, contractual terms and harm—not political labeling in place of legal analysis.
19. OpenAI’s Data Disputes and the DeepSeek Distillation Accusation Are Separate Issues
Lily’s rebuttal was that the US side is highly “hypocritical”: OpenAI scraped content from across the web during training, and outsiders had even questioned whether it downloaded YouTube videos without permission. Yet it now uses an industry-wide practice—distillation—to accuse a Chinese company of stealing technology.
张鹏 acknowledged that training data may involve copyright infringement, but separated the issues. Scraping text and images without permission is an input-side dispute; distillation uses data generated by frontier models and is a different issue that may also involve contractual relationships. The first cannot automatically cancel out the second.
The emotional basis for US companies is “free riding.” They paid the cost of training frontier models, and Chinese companies using outputs to shorten their development path can create a powerful sense of unfairness. But unfairness is not the same as satisfying the legal elements of intellectual-property infringement.
David Sacks has described distillation as intellectual-property infringement 3 times in succession and may push cloud providers to adopt KYC systems resembling bank anti-money-laundering controls to monitor and report distillation by Chinese companies. 张涛’s technical judgment remains that once a model is open to public use, truly preventing this is difficult.
20. DeepSeek’s First Regulatory Exposure Concerns App Privacy, Not Model Capability
Lily summarized the recent attention from the EU, US and Australia as likely to focus initially, in a narrow sense, on the app, data security and privacy. DeepSeek is young, and its investment and experience in product, operations and security are plainly below those of major incumbents; flaws are therefore unsurprising.
Her objection is that a startup’s first task is to develop the most innovative technology, and DeepSeek has already done that. Many websites and vendors around the world have similar privacy and security problems, yet DeepSeek’s breakout has made it the “tallest poppy.”
Globalization still requires local implementation. Lily noted that Sam had gone to Japan, where he formed a joint venture with SoftBank, and then to South Korea, where Kakao would handle ChatGPT’s local rollout. The logic resembles the “云上贵州” model of localization.
Lily is more worried about a TikTok-style outcome. If regulators simply define standards, a company can spend resources fixing problems; if the goal is to exclude Chinese companies, more compliance work may not change the result. TikTok has already spent enormous amounts of time, effort and money without eliminating its US political risk.
21. US AI Competition With China Now Covers Compute, Data, Capital and Talent
张鹏 traced the Biden administration’s main line back to compute: tightly restricting advanced GPUs, the semiconductor equipment needed to produce them, and upstream components and raw materials, with the aim of keeping China behind the US frontier.
Data rules seek to cut off flows of sensitive US data to China. Although the definitions were narrowed to protect the commercial interests of US companies, 张鹏 still expected the rules to take effect roughly a month after the program aired and to materially affect two-way data flows between the US and China.
Outbound investment screening has already framed 3 areas—advanced semiconductors, quantum computing and advanced AI. The core idea is that “US money cannot support China in developing capabilities that can compete with America’s frontier models.” Trump could theoretically revise the rules, but any change would have to be considered alongside tariffs and bilateral negotiations.
Talent could become the next layer of restrictions, including Chinese companies hiring local AI talent in the US and Chinese passport holders participating in the development of US frontier models. The policy discussions 张鹏 had heard involved visa and immigration tools, but there was no public decision at the time of the program.
22. DeepSeek Made the US Suspect That “Accelerate Yourself, Slow China” Had Failed
US policy treats whoever achieves AGI first as a matter affecting US-China strategic stability, almost at the level of nuclear weapons: use funding, industrial support and “Stargate” to make US companies run faster, while using export controls, capital restrictions and data limits to slow China.
DeepSeek’s shock is that Chinese models did not fall far behind as expected. Instead, they rapidly approached leading closed models and raised the possibility of overtaking them. 张鹏 believes this has created greater urgency in the US government, which is studying whether its existing tools worked and how to impose further limits.
The first 3 months after Trump’s inauguration, especially the first 100 days, were an important window for the top-level design of US AI competition with China. The Commerce Department, State Department, Treasury and White House National Security Council could conduct intensive reviews and submit the next package of restrictions to the president.
Potential actions include placing H20, which remained compliant at the time, under controls and pursuing more aggressive legislation. Josh Hawley’s proposed US-China AI Capability Decoupling Act aims at comprehensive separation across AI technology and intellectual property, R&D cooperation and capital flows; it is not the same thing as the self-media claim that a single download would incur a $100M fine.
23. Europe Looks More Like a Compliance Toll, While the US Treats Chinese Tech Companies as Security Threats
张鹏’s impression of Europe is that although the EU regulates aggressively, it generally remains more “case by case,” dealing with data, content and platform compliance rather than broadly elevating Chinese technology companies to the level of national security and industrial competition.
The high compliance requirements and huge fines under GDPR, the Digital Markets Act and the Digital Services Act have also been criticized as a disguised “regulatory tax.” Trump criticized the EU at Davos for using regulation to collect money from US companies. Businesses still have room to gain access through local remediation.
The US path is different. Since Trump’s first term, the rules have expanded from hardware into information and communications services, investment, two-way data flows, software and apps, increasingly resembling an effort to exclude Chinese companies from the US market.
AI applications are more foundational than social media, and US officials may see their security risks as greater. Products such as DeepSeek listed in US app stores could therefore become standalone regulatory targets rather than merely falling under existing privacy-law frameworks.
24. The Entrepreneurial Answer to Geopolitical Uncertainty Is to Operate First, Not Worry in Advance
张涛 believes Monica’s global market exposure is currently limited because it still relies mainly on OpenAI and Claude and does not train frontier models itself. The degree of confrontation and the rules have not yet been finalized: “There’s no point worrying now; we still have to see what action is ultimately taken.”
庄明浩 remains concerned about a TikTok-style outcome. If the EU merely introduces rules requiring compliance, a company can work to meet them; but if the other side intends to exclude it from the outset, even massive investments of compliance staff, effort and money may not change the result. He is “slightly pessimistic” about that possibility.
Their calm does not mean they deny the risk. It reflects the fact that entrepreneurs cannot use today’s emotions to solve rules that have not yet been defined; they can only wait for concrete action and continue responding.
25. Chatbot Traffic Is Not a Moat, and Mobile-Internet KPIs May Also Fail
Lily highlighted the contrast: 豆包 reached roughly 20M DAU after more than a year of promotion, Kimi never reached 10M, while DeepSeek R1 exceeded that scale within days. Users face almost no switching cost between models; traffic can surge overnight and change just as quickly.
庄明浩 believes this validates an earlier internal ByteDance view: the chatbot is not the ideal end state, but more like an industry-wide “second-best solution” that everyone can temporarily accept. Strategy should not focus exclusively on that form, and high DAU alone does not prove long-term value.
Citing January AI-product data, he said usage duration and retention remained weak across the industry. The DAU, MAU and paid-acquisition narratives familiar from mobile internet may not fit AI; DeepSeek itself may not even treat growing DAU or MAU as a KPI.
This is the fundamental pressure on existing model companies. If technical capability lags the leaders, users have no switching cost; if scale comes only from buying traffic, it is difficult to show that scale will translate into sustained usage and monetization.
26. Open-Source Models Have Caught Closed Models for the First Time, Giving the Application Layer More Combinatorial Space
张涛 noted that US opinion has already flipped into a joke: “Maybe in the end we’ll find that the biggest moat isn’t the model, but the wrapper.” In the past, even a strong open-source benchmark score still translated into an inferior real-world experience versus Claude or GPT-4. R1 was the first time the open-source world genuinely caught up.
Application companies no longer have to wait for OpenAI to open its latest API. After OpenAI released Deep Research, developers quickly used R1 with open-source Agent frameworks to reproduce similar experiences; some use cases previously dependent on closed models suddenly became testable.
R1 plus search is only the first example. The model can also be connected to RAG, documents and a range of tools. Different combinations, 张涛 believes, will create different application scenarios.
The industry still lacks prompt and product best practices for reasoning models, and users are experimenting with combining R1 with Claude, GPT and Cursor. 庄明浩 used OpenAI’s L1-to-L5 framework to describe the next step: chatbot is L1, reasoning models are L2, and the industry should move toward L3 Agent, where models actually complete tasks.
27. China’s “Six Little Dragons” Must Rechoose Their Technology, Open-Source and Commercialization Paths
庄明浩 believes companies that shifted direction early face relatively limited disruption. In his summary, 一零一物已 handed its NVIDIA GPUs and training team to Alibaba and now focuses on solution implementation, while 百川 concentrates on healthcare. Neither is still making a direct bet on winning the general-purpose model race.
智谱 may be less affected because it formed a relatively clear technical route and execution rhythm early on. 阶跃星辰 was founded later and remains less understood externally, so 庄明浩 declined to force a judgment.
The questions are sharper for Kimi and MiniMax. Kimi has released a model that displays its reasoning process and delivers a good user experience, but remains closed source; it must decide whether to follow the open-source path and how to continue its existing commercialization strategy. MiniMax CEO 闫俊杰 acknowledged in an interview that it may have been better to open source from the beginning.
For both companies, some of their strategic investment over the past 2 years may amount to “wasted time, wasted money and wasted people.” If they still want to climb the general-purpose technology peak, they must contend with both DeepSeek and Alibaba’s Tongyi while rebalancing open source, closed source, fundraising and revenue.
28. DeepSeek’s Structure Lets It Defer the Commercialization Question for Now
Open source does not imply a single economic model. Some projects release weights but require commercial users to buy a license; parts of Stable Diffusion and Tongyi Qianwen use arrangements of this kind. DeepSeek uses the MIT license, however, so third parties can deploy and provide services without buying a commercial license from it.
张涛 admitted that he could not yet see how DeepSeek would make money directly. 庄明浩’s key point was that DeepSeek has no outside investors, while its parent 幻方 has deep financial resources. For at least the next 2 or 3 years, it can invest with the freedom of a research institution, without immediately answering to financiers or commercial commitments as OpenAI must.
张涛’s underlying belief is: “As long as you keep creating value, as long as you are doing genuinely valuable work, that value will eventually be monetized in some form.” If DeepSeek creates AI capabilities of major value, monetization need not necessarily appear as cash revenue alone.
庄明浩 called it a structural anomaly at the table: “An anomaly has appeared at the table.” Its origins, resources and objectives differ from those of the other players. It currently seems best suited to long-term technical investment, but whether it ultimately wins remains, as the program put it, “something no one knows.”
29. R1 Reversed the Debate Between “Commercial Victory” and “Technological Idealism”
Lily recalled 3 shifts in public opinion. In March and April last year, 朱啸虎 said he “couldn’t look at or invest in a single foundation-model company,” drawing heavy criticism. When companies such as 一零一物 made adjustments and the industry began emphasizing revenue, the market treated it as a victory for 朱啸虎.
Less than a month later, R1 won global users with technical capability that in some respects matched the industry leaders, and public opinion once again called it a victory for “technological idealists.” 豆包, Kimi and others had spent heavily on promotion without achieving anything comparable, leading Lily to conclude that “technology is the primary productive force.”
庄明浩’s reminder was that investing in frontier foundation models is inherently a high-risk bet. Participants cannot deny the risk only after an anomaly appears at the table. DeepSeek’s strategy has forced everyone to place new bets, but the current lead structure does not guarantee the final outcome.
30. Enterprises Should Stop Training Their Own Pre-Training Models, but Cannot Keep Waiting on the Sidelines
In response to legal firms and traditional companies asking whether investment would quickly become obsolete, 张涛’s answer was direct: “Stop investing in pre-training.” Base models will become an upstream commodity over time, so companies can buy or use existing models instead.
Most vertical use cases do not even need sophisticated post-training; getting RAG right may be enough to start. Mature post-training methods can be added where necessary. Private deployment also does not require developing a base model from scratch: an MIT-licensed R1 can run in a company’s own environment.
庄明浩 relayed a view from a company’s annual meeting: many traditional-industry CEOs still want to “let the bullets fly for a while,” but given the pace of change over the past 2 years, those who still do not “physically engage in depth” now may have no opportunity later.
The investment priority is not betting on whichever version is strongest today, but using and participating in AI businesses as early as possible. The technology will continue to change, but staying entirely outside the process will not reduce uncertainty.
31. The Final 2 Rumors Both Conflate “Can Run” With “Is the Real R1”
庄明浩 specifically denied a widely circulated story that, when DeepSeek’s service came under pressure, domestic security vendors, Huawei Cloud, Tencent and others jointly provided support. The story was “completely false,” a wolf-warrior-style mutual-aid narrative designed to ease users’ anxiety when they could not access the service.
The rumor 张涛 repeatedly debunked concerned “running R1 locally on a phone or computer.” Most tutorials actually run DeepSeek-R1-Distill-Qwen-32B, 7B or even 1.5B, or the corresponding Llama versions. From the naming syntax, the head noun remains Qwen or Llama; R1 merely provided distillation guidance.
These small models use neither V3 as their base model nor R1’s reinforcement-learning process. Their publication mainly demonstrated that SFT data generated by R1 could improve other models. They are nowhere near the full-size R1 in capability, and successfully starting a local model is not the same as replicating R1.
张涛 does not oppose using the hype to learn local deployment. If the goal is to understand model execution, memory requirements and the toolchain, the tutorials still have educational value. But users should not pay “full-strength R1” prices for them, nor use them to judge the capabilities of the real R1.