Vol.83 A 25-Slide PPT Captures China’s AI “All-Star Game”
Summary
- One important conclusion from the forum was that the ChatGPT battle seems to be over, and the next one will shift to Agents that actually do things. 唐杰 sees the next battle as solving real-world problems through Agents; over the past year, Zhipu has also pushed into Coding and Agent. 姚顺禹 goes further, arguing that even if model R&D stopped entirely today, deploying existing capabilities across To B could still offer 10–100x room for growth. The key tests of commercial value will be task duration, deployment scale, and whether “problem value, operating cost, and iteration speed” can be balanced.
- Models and applications are splitting along To C and To B into two distinct logics: the former is moving toward tightly coupled All in One products, while the latter is becoming more modular and more dependent on strong models. 姚顺禹 believes ordinary users may not feel a major difference when asking ChatGPT questions today versus a year ago, while enterprises are continuously embedding stronger models into their operations, with DingTalk as a prime example. Tencent’s To C focus is complex context and memory; on To B, simply transforming its many internal scenarios is “already big enough.”
- The most cautious estimate at the forum for China producing the world’s leading AI company within 3–5 years came from 林俊旸: “Even 20% may be optimistic.” He estimates that China and the US differ by one to two orders of magnitude in compute, while leading US suppliers have already directed more R&D compute and capabilities toward the next paradigm. 姚顺禹 is relatively more optimistic, citing China’s progress in power, infrastructure, and talent, though lithography, To B delivery, and willingness to pay remain constraints. 林俊旸 even worries that the gap may still be widening.
- Continual learning is already a Silicon Valley consensus, but 姚顺禹 believes it will not arrive as a product suddenly declaring victory; it will be a gradual process. Personalized chat and Coding tools that adapt to a company’s coding habits—even the fact that “95% of Claude Code’s own code was written by AI”—could be early forms in constrained settings. 林俊旸 also warns that when the environment, not just the user’s prompt, can stimulate a model to act, more security problems will follow.
- In 唐杰’s view, multimodality instead made 2025 a “year of disappointment,” implying that 2026 could bring a more aggressive catch-up effort and heavier investment. Qwen now places advanced reasoning, top-tier coding capabilities, and native full multimodality on the same priority list, targeting “three in, three out,” with multimodal coverage for both inputs and outputs. Its image-editing workflow overlays two images to check whether unedited regions drift—a harder engineering metric than demo quality—and the latest 2.5.1 release has already materially improved the problem.
- Kimi’s differentiation is not a pile of features, but higher token ROI, longer context and memory, and the creator’s taste embedded in model generation. 杨植麟 says that asking “what values should a good AI pursue?” gives the model a taste that distinguishes it from others: “It is not a commodity.” He disclosed that K3 is coming soon, explaining his choice with the view that pessimism means stagnation, while optimism, even with its risks, represents unlimited possibility.
- 张钹’s AGI yardstick is not better conversation, but being “executable and verifiable”: multimodal understanding, online adaptation, long-term planning, reflection and cognition, and cross-task generalization are all indispensable. The corresponding paths span embodied interaction, grounding in evidence, knowledge alignment, tool execution, and constraint governance; the most important object of governance, he argues, “is not machines, but humanity.” For enterprises, this shifts competition beyond model capability toward whether knowledge, ethics, and applications can become reusable and broadly accessible infrastructure like water and electricity.
Deep dive
1. A High-Caliber Lineup Recalibrates the Roadmap for China’s Open-Source Foundation Models
AGI Next, held on January 10, was organized by Tsinghua University’s Beijing Key Laboratory for Foundation Models and convened by 唐杰. 杨植麟 and 林俊旸 spoke in succession; 杨强, 唐杰, 林俊旸, and 姚顺禹 joined a roundtable moderated by 广密; 91-year-old 张钹 delivered the closing summary. Four academicians were in attendance, and the opening guests included the president of Tsinghua University and the party secretary of Haidian District.
庄明浩 called the lineup China’s “starting lineup of all-stars” in foundation models, while the audience joked that it was an “exchange meeting for Nascent-Soul powerhouses.” Apart from MiniMax, which had just gone public and was probably busy with its annual meeting that evening; DeepSeek, which almost never appears or speaks publicly; and ByteDance’s Seed team, focused mainly on closed-source models, the industry’s core forces were highly concentrated.
He fed the event transcript, photos, and personal notes into NotebookLM, which first generated 3–4 PPTs before he manually consolidated them into 25 pages. It was his first time using an AI-generated PPT as the presentation material for the main program. “NotebookLM did most of the work; I did a small part of the work” also became a footnote to AI’s entry into the content-production workflow.
2. 唐杰 Points to a Risk Bigger Than Open-Source Gains: “The Gap May Still Be Widening”
唐杰’s review of the evolution chain ran from small models and large models through Transformer to today’s SOTA; capabilities have progressed from language to Coding and Agent. In 2025, RLVR—reinforcement learning with verifiable rewards—became a mainstream direction. Zhipu has also worked on Coding and AutoGLM.
His most sober warning was: “The gap may still be widening.” US foundation models are evolving more heavily along the closed-source path, and China cannot become complacent simply because of its open-source progress.
唐杰 still defined the first task for 2026 as exploring the upper bound of intelligence: scaling along known paths so foundation models can enter Coding, Agent, and eventually the physical world, while searching for new paradigms beyond o-series reasoning.
The technical agenda includes new model architectures, ultra-long context, more efficient knowledge compression, memory mechanisms and continual learning, multimodal fusion, complex-task Agents, robotics, and AI for Science. He considers “model autonomous learning” the key to AI surpassing humans and moving toward superintelligence, making it the most frequently repeated keyword of the day.
3. 杨植麟 Bets Kimi on Token Efficiency, Long Content, and Model Taste
With “return to first principles and push the frontier of AGI” as the central framing, 杨植麟 reduced Kimi’s training roadmap to two priorities: improving token efficiency and handling ultra-long content, context, and memory. He said the current improvement in AI intelligence is fundamentally a shift “from energy to intelligence,” which is why he places particular emphasis on token ROI.
He believes model builders inevitably embed their own taste into the model’s creative process: “What values should a good AI pursue? That taste distinguishes one model from another. It is not a commodity.”
On how to live with AI, he said: “I choose optimism, because pessimism means stagnation, while optimism, even with its risks, represents unlimited possibility.” He pointed AGI’s ultimate vision toward unknown problems in science, health, and energy, adding that the clear focus today is K2, with K3 coming soon.
4. Qwen Puts 2026’s Decisive Edge on Native Full Multimodality
林俊旸 first recommended that users experience Qwen’s various capabilities through chat.qwen.ai rather than the app. Qwen spans many sizes, from 1.8B, 7B, and 72B to possible intermediate models such as 3B and 4B, each added in response to customer demand.
Its current core is advanced reasoning and context, industry-leading coding capability, and “genuinely native full multimodal capability.” 林俊旸 expects Qwen to achieve “three in, three out” in 2026, with multiple modalities covered in both inputs and outputs. His value judgment was direct: “If you are not building foundation models for all humanity, then don’t build them.”
The image-generation example was not a single beautiful image, but a possible 12-panel “day in the life of a person,” with each panel simultaneously generating a fixed time, corresponding text, and the person’s action. 林俊旸 used it to show that image models are gaining comprehension capabilities similar to Nano Banana, rather than merely synthesizing visual elements from keywords.
The harder test came from the community: ask a female model to lower her right hand, then overlay the images before and after editing. If the clothing, body, or face also becomes blurred beyond the arm, the unedited regions have drifted. Qwen iterated accordingly, and the latest 2.5.1 version showed a major improvement in drift, a detail that led 庄明浩 to expect further progress in multimodality and full multimodality.
5. The Foundation-Model Race Is Splitting into Two Logics: To C and To B
庄明浩’s biggest observation about 姚顺禹’s public appearance as Tencent’s top AI executive was that he split all four questions into To B and To C before answering them. That strong sense of boundaries may reflect his interactions, since joining Tencent, with management, business units, and China’s To C internet ecosystem.
姚顺禹 believes a To C user asking ChatGPT a question today may not feel a huge difference from the answer received a year ago, and users do not necessarily need the strongest model at all times. To B, by contrast, is using stronger models and embedding them more deeply into operations, with DingTalk as a prime example. The application layer is splitting accordingly: To C is tightly coupled and moving toward All in One, while To B is becoming much more modular.
When pressed on what Tencent is actually betting on, he again answered from both ends: To C is centered on complex context and memory; To B is harder, but Tencent already has a large number of scenarios and businesses, and simply applying model capabilities to internal productivity gains is “already big enough.”
Other guests offered different angles. 林俊旸 emphasized starting from real user needs and argued that companies may not have fixed genes: “Generation after generation may shape these companies.” 杨强 said academia is catching up after seeing industry move at speed. 唐杰 judged that “the ChatGPT battle seems to be over”; the next battle is Agent doing things.
6. Continual Learning Is a Consensus, but More a Gradual Shift Than a Breakthrough
姚顺禹 said Silicon Valley has reached a consensus on autonomous learning, but the concepts of online learning and continual learning remain fuzzy, and different people may not be discussing the same thing. He leans toward seeing this not as a single methodology or a brand-new research paradigm, but as a problem of data, tasks, and the corresponding reward functions.
Personalized chat tools will gradually fit users’ styles, while Coding tools will increasingly understand a company’s coding habits. Does that already count as continual learning? He offered two boundary cases: ChatGPT fitting conversation style to user data, and “95% of Claude Code’s own code being written by AI.” In his view, continual or autonomous learning has already occurred in some specific settings, but remains constrained by the scenario.
林俊旸 said domestic model companies have not invested as aggressively in RL compute, while overseas players such as xAI are putting large amounts of compute into reinforcement learning. In AI for Science, the model finds materials, defines questions, obtains results through experiments, and produces a report; that process itself appears to be a form of learning. But when the environment can actively stimulate the model, more security problems will also arise.
唐杰 was relatively optimistic, saying 2026 could bring paradigm innovation. The efficiency bottleneck in scaling is already visible, and academia is beginning to build experience, even though universities may still have an order of magnitude less compute than industry. For model companies, the real decision is how to optimize the return on pretraining, post-training, new data, and new paradigms within a fixed compute budget.
7. Agents’ Economic Value Will Be Realized First in To B, Then Extend into the Physical World
姚顺禹 believes the To B curve will keep rising in plain sight: even if all model R&D stopped immediately, simply continuing to penetrate and deploy current capabilities could still offer 10–100x room for growth. In 2026, the expectation is that Agents will execute tasks for longer periods and handle more deployment work.
The deeper the collaboration between people and Agents, the more important education becomes, because learning how to bring out human capabilities and help Agents improve will be critical. 庄明浩 linked this to the “Fermi level” described by 田渊栋: how good and how powerful a person can make AI may determine that person’s future value.
林俊旸 believes Agents currently run mainly in computer environments, but in 2026 they will move more into non-computer environments, naturally extending into robotics and embodied intelligence. 杨强 divided the field into four quadrants based on whether goals and plans are generated by humans or machines, arguing that many of the positions visible today remain in the early first quadrant, while the other quadrants remain worth anticipating.
唐杰 reduced application success to three variables that must hold simultaneously: “the value of solving the problem, the cost of running it, and the speed of iteration.” Whether an Agent can create economic value depends not only on its capability ceiling, but also on whether each run is worth the cost and whether the system can iterate quickly.
8. China’s Odds of Producing the World’s Leading AI Company Fall to “Even 20% Is Optimistic”
姚顺禹 was relatively optimistic about the next 3–5 years. On power, infrastructure, and talent, China is “keeping very close and getting closer.” He still acknowledged practical constraints including lithography, To B market delivery, and willingness to pay, but believes these problems could still turn around. He also expects more risk-taking talent capable of innovation to emerge.
林俊旸 was markedly more cautious, estimating that China trails the US by one to two orders of magnitude in compute, while leading US suppliers have already directed more R&D compute and capabilities toward next-generation paradigms. Alibaba’s chip chief once asked, 3 years in advance, whether Transformer and multimodality would still be the future, because the tapeout cycle is 3 years. 林俊旸’s answer at the time was: “I don’t even know whether I’ll still be at Alibaba 3 years from now.”
When 广密 demanded a specific probability, 林俊旸 said “even 20% may be optimistic” and worried that the gap might still be widening. 唐杰 gave no number, instead summarizing the necessary conditions as a group of smart, risk-taking people; a better business environment; and more people willing to persist over the long term with “dumb” work.
9. 张钹 Places Verifiable AGI’s Endpoint in Governance and Entrepreneurial Responsibility
张钹 called for moving from language models to Agents, and from language to behavior, so that systems can complete complex tasks in complex environments. His “executable and verifiable AGI” requires five capabilities: spatiotemporally consistent multimodal understanding and grounding, controllable online learning and adaptation, verifiable reasoning and long-term planning and execution, calibrated reflection and cognition, and strong cross-task generalization.
The corresponding research paths include multimodality, embodiment and interaction, retrieval and grounding in evidence, structured knowledge alignment, tool execution and deployment, and alignment and constraints. These paths may map onto the capabilities above.
He divided AI agency into 3 layers: from function to an agent of action, then to a subject of norms and responsibility, and ultimately to a subject that may possess consciousness and experience. Working backward from the end state, the urgent priorities remain alignment and governance—and “the most important governance is not governance of machines, but governance of humanity.”
张钹 said AI companies will no longer provide only products and services, but will make knowledge, ethics, and applications into reusable tools, ultimately turning artificial intelligence into a general-purpose technology like water and electricity. Entrepreneurs are therefore in an “honorable and sacred profession,” and must also assume responsibility for governance, broad access, and sustainable growth. At 91, he still went to the lounge during the event to revise the final PPT. 庄明浩 described himself as “just a recorder” and closed the account with the thought that “doing the real thing may be the only way out.”