Pioneers Insight Method Research Author
20. From 2024 to 2025: Review, Assessment, and Outlook on the Large-Model Wave
Back to Episodes

20. From 2024 to 2025: Review, Assessment, and Outlook on the Large-Model Wave

Summary

  • 2024’s central theme was the industry’s shift from exuberance to a cooling-off period, and from benchmark competition to calibrating technical paths against user value and a viable commercial loop. 李子玄 called it “start with the end in mind” (以终为始); 樊家睿 distilled the year into “multimodality, cognition, and foresight,” while 李杨 called it a “return to rationality.” AI-native products and Super Apps will not emerge overnight; opportunities are being created by new solutions to old needs.

  • Kuaishou’s case shows that big tech’s advantage lies not only in resources, but also in embedding models into existing data, traffic, and monetization loops. “Upgrade understanding, improve interaction, and explore generation” maps respectively to recommendation and governance, AI Xiaokuai with MAU in the tens of millions, and an LLM-plus-digital-human livestreaming product that drives roughly RMB30M in daily ad spend. Cross-functional execution still requires a single objective, concentrated resources, and organizational resilience.

  • Kling AI chose to charge early, essentially using payment behavior to validate PMF while trying to cover expensive training and inference costs. In roughly 5 months from July, it iterated nearly 20 times; users had reached 5M, and first-month billings topped RMB10M. 李杨 put the logic bluntly: “If we explored this like a traditional internet company, we wouldn’t even have a chance to recover cash—we might simply be killed off” (如果我们像传统互联网那样去探索,我们连回血的机会都没有,可能就被干废了). Its subsequent overseas expansion also produced positive feedback on both targets and revenue.

  • Big tech has traffic and deep pockets, but startups’ defensible value lies in technical foresight, understanding model boundaries, and owning a differentiated niche. 樊家睿 said ShengShu Technology had committed to a Diffusion Transformer fusion architecture since 2022, or perhaps earlier, pursuing a general multimodal system capable of arbitrary inputs and outputs. 李子玄, meanwhile, cautioned that “Doubao is genuinely too strong,” but only for one category of mass-market demand; J people, P people, F people, and T people can still support differentiated markets.

  • Multimodal competition has moved beyond whether generation is usable to whether consistency, controllability, and speed can unlock new product forms. ShengShu Technology’s progression ran from facial consistency in July to subject consistency in September, then multi-subject, multi-angle, and multi-scene consistency in November. Its view is that Chinese teams are moving from following to “being native to this time and place,” though film and television still remain far from true screen-grade industrial deployment.

  • Zhipu exposed the application’s key unresolved problem: demonstrating capability is not the same as finding a stable, paying user story. Video calls and AutoGLM ordering takeout are still at the zero-to-one stage. The same action might mean repeatedly ordering Kung Pao chicken from Hehe Gong, or spending 30 minutes selecting zero-calorie sugar; meanwhile, B2B token prices could fall to 1/100, 1/1,000, or even 1/10,000 of their former level. Usage growth alone cannot make up the gap; models must take on more complex, higher-value tasks.

  • The 2025 outlook is for richer application soil, with a multimodal breakthrough in efficiency and generality potentially triggering a black swan. 李杨 expects technological discontinuities and a proliferation of vertical applications, but does not assume the black swan will be positive or negative. 李子玄 urged companies to build moats while accounting for human solutions and other substitutes. 樊家睿 offered a more concrete threshold: Vidu takes roughly 20 seconds to generate a 4-second video; only when generation is no slower than the clip itself will real-time interaction, infinite generation, and entirely new content categories truly open up.

Deep dive

1. 2024 Entered a Cooling-Off Period—but Returned from Benchmarks to Outcomes

  • 李子玄 used “start with the end in mind” (以终为始) to describe the product shift: the objective cannot stop at model leaderboards, but must become a user metric the team genuinely optimizes. Once the objective becomes the metric, the objective itself may cease to matter.

  • 李子玄 cited Perplexity as an example: its 6-month retention rate is 45%, with a goal of learning from Duolingo’s 50%; it is also trying to increase the number of searches users make in their first conversation. Behavioral metrics, rather than a single-minded race on benchmarks, are being used to measure user mindset and experience.

  • 樊家睿’s keywords were “multimodality, cognition, and foresight.” His core view is that Chinese teams have begun choosing their own technical directions, including consistency and speed; “they are native to this time and place,” signaling that industry autonomy and technical leadership are advancing together.

  • 李杨 called 2024 a “return to rationality”: language models are moving toward deeper reasoning, while visual generation is iterating faster. On the product side, the industry is no longer fixated on immediately discovering an AI-native product or Super App, but is instead “looking for new opportunities inside old needs.”

2. AI Capability Demos Still Leave Dirty Work Before Real User Stories

  • Zhipu Qingyan’s video-call feature can identify scenes outside the camera’s frame and interact with the user. 李子玄 called it “a landmark feature,” but acknowledged that it remains at the zero-to-one stage: why users would use it and what need the model actually solves have not been fully defined. Use cases such as solving problems or practicing English have been observed, but remain incomplete.

  • AutoGLM’s takeout-ordering feature exposed enormous variation within the same scenario: one user may only want to reorder Kung Pao chicken from Hehe Gong, while another may browse for 30 minutes and insist on zero-calorie sugar. If AutoGLM ignores those preferences, its supposed automation can instead feel as if it “doesn’t really understand me.”

  • 张海庚’s point of agreement was that getting technology through the last mile to users is not romantic; it is full of “dirty work.” Teams must map complete user stories and compare substitute solutions, but also leave the area around Tsinghua University to understand how people who do not use an iPhone, do not use a Mac, or may not even know Chrome perceive these products.

3. New Demand and Old Scenarios Have 3 Different Relationships

  • 樊家睿 divided the market into 3 categories. The first is new demand exploding out of old scenarios at a scale far beyond expectations: the AI content users create on traffic platforms has already exceeded what traditional technology could deliver.

  • The second, represented by games and anime, is new demand that is temporarily attached to old scenarios but could evolve into different product forms through real-time interaction and a range of content-generation capabilities. The problem is that current technology still cannot cover the industry’s requirements end to end.

  • The third includes short dramas and film and television, where old scenarios demand more realism and controllability than current models can provide. 樊家睿 did not frame the gap as a permanent constraint: as foundation models continue to evolve, new demand could eventually break free from the original production model.

4. Kuaishou Embedded Large Models in Its Core Business Before Exploring External Products

  • 李杨 emphasized that Kuaishou is not taking open-source models and fine-tuning them a second time; it is training “from scratch.” Beyond language models, it is also working on image, audio, TTS, text-to-image through Ketu, and video generation.

  • Kuaishou’s internal application strategy was compressed into 3 phrases: “upgrade understanding, improve interaction, and explore generation.” Upgrading understanding uses short-video, livestream, and comment data to improve recommendation and platform governance, and Kuaishou has already seen gains in both areas.

  • “Improve interaction” takes a deliberately restrained approach: AI Xiaokuai responds only when users actively @ it and ask a question in the comments. Its role is to supplement the information in a video, not to occupy the community with bots. Even so, its MAU has reached the tens of millions.

  • “Explore generation” is already linked to revenue: combining an LLM with digital-human livestreaming capabilities generates roughly RMB30M in daily ad spend for Kuaishou, and 李杨 stressed that this is an average.

5. Kling’s Rapid Iteration Came from User Feedback and Feedback from Action

  • After Kuaishou officially released Kling AI on June 6, its reach and use cases exceeded the team’s expectations. Its 2 operating principles are “learning from users” and “learning from acting”: first use user feedback to decide what to build, then use product experiments to validate priorities and feed the results back into model development.

  • In roughly 5 months from July, Kling iterated nearly 20 times, close to a weekly release cadence. 李杨 emphasized that this was not an aimless chase for version numbers, but a continuous effort to answer which feature fits which scenario and what the scenario’s pain points are, then return those answers to model R&D.

6. The Elephant Can Dance, but Coordination Drag Must Be Compressed

  • Asked whether large companies are inherently slow, 李杨 acknowledged that coordination across departments and roles is more difficult. Kuaishou’s solution has 2 parts: narrow the phase objective to a single priority, and integrate underlying resources with horizontal teams so everyone repeatedly aligns around the same deliverable.

  • Kling did not begin only after Sora was released. Kuaishou had been investing in video generation since late 2023; Sora’s lesson was to switch quickly to the DiT architecture. The team at the time considered making the product “usable” within 1 year an optimistic estimate, but still directed substantial compute and data resources toward it.

  • 卫诗婕 asked whether Kuaishou was using large models to restore morale. 李杨 replied that “after Kling AI came out, the whole company should have been buzzing.” His 2-track metaphor was “climbing a mountain while sailing”: use AI to overtake the short-video core business on the inside, while incubating a new high-value product on the outside.

7. Kling Narrowed Its Focus 3 Times Through a Materials Wedge, Monetization, and Overseas Expansion

  • 李杨 said Kling is not currently replacing the entire video industry. It is first entering the value chain for production assets: some shots that previously required live-action filming or licensed libraries can now be generated by the model. The model is imperfect, but the entry point is already viable.

  • Kling had announced a user base of 5M and began monetization as soon as it formally opened to the public; first-month billings exceeded RMB10M. The decision to charge quickly was criticized as premature, but the team believed the cost of training and inference made it impossible to copy the internet playbook of offering products free first to exchange for scale.

  • His formulation left no room for hedging: “We wouldn’t even have a chance to recover cash; we might simply be killed off” (我们连回血的机会都没有,可能就被干废了). Charging was not only an effort to recover costs, but also a way to test whether a scenario deserved continued investment and whether PMF existed.

  • The team then expanded overseas, where the product received positive feedback on both targets and revenue. Overseas creators often use multiple tools simultaneously, leaving room for products with stronger performance in a single capability.

8. A Startup’s Defense Is Route Awareness, Not a Traffic War with Big Tech

  • 樊家睿 acknowledged the inherent traffic advantage of internet giants, but argued that startups can use technical understanding, foresight into the end-state of products, and feedback loops among technology, users, and markets to “eat a very thick slice” of the value chain.

  • ShengShu Technology said its team had published a paper on a Diffusion Transformer fusion architecture as early as September 2022, before the Sora team’s work on DiT. Its route has consistently targeted a simple, efficient, general-purpose multimodal foundation rather than switching directions with each market hot spot.

  • His key distinction is that many systems now labeled “multimodal” are still just stacks of parameters from text, video, and audio models. A truly general foundation should support arbitrary modalities as both inputs and outputs. At a stage where “model boundaries determine product boundaries,” understanding those boundaries and continuously pushing them outward is a major startup advantage.

9. Doubao Is Powerful, but Has Not Eliminated the Niches for Chatbots and Agents

  • 李子玄 was direct: “Doubao is genuinely too strong—there’s no question about it” (豆包确实太强了,这一点毋庸置疑). But Doubao and Kimi already have different user positioning, while their retention rates are currently similar—evidence that most products are still merely calling models and that personalization has not truly emerged.

  • 李子玄 used “J people, P people, F people, and T people” to show that demand will not become completely homogeneous. Doubao may satisfy the largest mass-market segment, but that does not mean other groups lack a market; Doubao itself also faces retention risk, and current results do not determine the final winner.

  • Facing handset makers as well as major players such as Amap and Baidu Maps, AutoGLM chose a visual approach to improve generalization: even if the model has never seen a particular app, it should be able to recognize the interface and complete the operation, avoiding the high maintenance costs that traditional RPA incurs when devices, APIs, and buttons change.

  • 李子玄 also kept his skepticism intact: “What we’re saying today may still be only a hypothesis.” Handset makers could improve generalization as well, but their AI capabilities often serve new-phone sales. AutoGLM may instead be better suited to penetrating the installed-base market and dividing up ecosystem roles through partnerships, rather than treating every handset maker as an adversary.

10. Apps Will Not Simply Disappear; Aggregation and Niche Layers Will Grow in Parallel

  • Responding to the prediction that there will be fewer applications in the future, 李杨 said 2 earlier claims had been conflated. Large apps with more DAU may become fewer, while plugins and small applications built by individual developers may become more numerous; both trends can hold simultaneously.

  • AutoGLM is better understood as a scarce orchestration layer, not a replacement for loyalty and membership systems. Users will not abandon Meituan simply because Ele.me adds automatic ordering, nor will they readily give up their Didi memberships. AutoGLM should complete tasks across apps, not rebuild every app.

  • Based on conversations overseas, 李杨 explicitly disagreed with the domestic preference for “big and complete” products. Overseas users are more inclined toward finely tuned services in narrow fields, so he “doesn’t really agree that there will be fewer applications in the future.”

11. Vidu Is Using Consistency to Connect Technical Breakthroughs Back to Industrial Demand

  • ShengShu Technology said that despite relatively weak operational growth, Vidu became one of the world’s largest video-generation products by user base, driven mainly by organic traffic. It initially targeted professional and semi-professional creators, then found that ordinary users’ willingness both to embrace AI and to pay for it was far higher than expected.

  • The first B2B opportunity is internet entertainment content that traditional technology cannot produce, such as interactions among multiple IPs. The second is new workflows for games and anime. The third is film and television, where full-length production does not yet meet requirements, but promotion, marketing materials, and production assets are already supporting mature partnerships.

  • A representative case is the AIGC short film for Venom: The Last Dance, produced in partnership with Sony China. 樊家睿 called it the world’s first AIGC-generated film short to receive authorization from a major IP, as well as the first commercial collaboration between a Chinese large-model company and a top international IP. Moving the project forward required simultaneous work on compliance, legal, and IP issues.

  • Industrial feedback has driven the technical roadmap: facial consistency in July, subject consistency in September, and multi-subject plus multi-angle and multi-scene consistency in November. Extending from a single image, without separate fine-tuning, to people, animals, products, tri-fold phones, Tesla vehicles, and virtual characters is a major direction of the evolution.

12. Sora Brought Resource Validation, Not a Strategic Pivot at ShengShu Technology

  • Responding to claims that the company only raised video to its top priority after Sora appeared, 樊家睿 corrected the record directly: by late 2023, the team already had a video prototype based on its fusion architecture, with high frame-by-frame quality and the ability to generate 4-second videos.

  • The real constraint was that the market had been dominated by the idea that “Transformer can compress everything.” Teams committed to fusion architectures struggled to secure enough resources. After Sora was released, the outside world looked back and validated the route, and resources began to concentrate there.

  • 樊家睿 therefore still thanked Sora for “releasing at the right time,” but rejected a narrative of Chinese progress as mere following. He said the team’s advances—from consistency to the emergent intelligence, contextual understanding, and associative capabilities of vision models—were “native to this time and place,” rather than a response to technological leadership emerging in the United States.

13. B2B Price Cuts Forced Zhipu to Recalculate Its Revenue Formula and Charging Timeline

  • 李子玄 breaks the business down as revenue minus costs, then breaks revenue down into P multiplied by Q. The old assumption was that token call volume Q would rise while price P stayed relatively stable. But price could fall to 1/100, 1/1,000, or even 1/10,000 of its former level, and usage cannot grow 10,000-fold in tandem.

  • 李子玄 sees the opportunity not in generating more simple text, but in consuming more tokens to complete higher-value tasks, such as deeper reasoning or having AutoGLM plan an entire operating path. Only these tasks may be able to inject new energy into B2B.

  • Zhipu is looking at network effects and lightweight growth simultaneously, rather than mechanically separating B2B from C2C. An open platform can serve creators on one side and empower partners through APIs on the other. The key question is what moat is being built, not which side of the revenue ledger the business belongs on.

  • C2C monetization remains cautious: models close to the first tier remain free, while a potentially highest-quality model is offered for a limited number of uses. Memberships are mainly being used to test demand and segment users. 李子玄 believes the company should first watch retention and get close to PMF before charging; e-commerce, a marketplace, and partnerships with other apps as an entry point or channel remain unproven.

14. The 2025 Watershed: Can Moats and Real-Time Generation Arrive Together?

  • 李杨 expects a technological “black swan,” while making clear that it may not be good or bad. He is more confident that applications will proliferate, because technology is still advancing, users are forming their understanding and habits, and vertical teams are beginning to focus and receive feedback. 张海庚 added that open-source progress and broader access to technology will also be critical to deepening the application soil.

  • 李子玄 cautioned that “a product is only an outcome; it is not a capability” (产品只是一个结果,它不是一个能力). Doubao’s Douyin distribution channel is difficult to replace, while the Minimax team’s understanding of US Gen Z and Gen Alpha may also be deeper. Durable advantages still come from company culture, principles, talent incentives, resources, and relationships.

  • 李子玄 said companies must also count non-technical substitutes. Some law firms in tier-4 and tier-5 cities hire paid interns; if they adopt AI or eliminate interns, they may instead lose the RMB200 they could have earned. “The complexity of human nature” means AI does not automatically create a purchasing motive through cost reduction and efficiency gains.

  • 樊家睿 summarized 2025 as multimodality becoming more “efficient and general.” Audio-video and 3D-scene video, more forms of consistency, and longer windows for contextual understanding and association will continue to converge. Vidu takes roughly 20 seconds to generate a 4-second video; if generation can be completed in no more than the video’s duration, real-time interaction, infinite generation, and entirely new categories of content consumption may become possible.