Pioneers Insight Method Research Author
98: 李开复 on Part of 01.AI Joining Alibaba: Only Big Tech Can Pursue Frontier-Scale Models
Back to Episodes

98: 李开复 on Part of 01.AI Joining Alibaba: Only Big Tech Can Pursue Frontier-Scale Models

Summary

  • 01.AI was not “acquired by Alibaba”; it voluntarily exited the frontier-model arms race, forming a joint lab with Alibaba Cloud in which most of its pretraining and AI infra teams joined and became Alibaba employees, while 01.AI retained a smaller team for models and applications. 李开复 explicitly said the company is “giving up on” frontier-scale-model pretraining, though Yi-Lightning V2 may still be pretrained once or twice a year, depending on returns. The new division of labor: Alibaba trains the teacher model, while 01.AI builds a student model that is “small enough, fast enough, cheap enough, and capable enough.”
  • 李开复 believes training-time Scaling Law has entered a zone of diminishing returns, making frontier-scale models increasingly suitable only for companies genuinely pursuing AGI and able to absorb enormous costs—typically, in his view, Big Tech. His analogy: “Going from one GPU to 10 might deliver the value of 9.5 GPUs; going from 100,000 GPUs to 1 million might deliver only the value of 130,000.” In China, he is currently fairly certain Alibaba and ByteDance can continue; startups need an exceptionally long, hard-to-replicate advantage, otherwise competing on how much they can burn is unlikely to work.
  • 01.AI is betting on fast models not only because inference is cheaper, but because test-time Scaling Law magnifies the commercial value of speed. Yi-Lightning ranked sixth globally at the time, ran several times faster than Yi-Large, and cost roughly one-thirtieth as much as GPT-4o. Producing tokens five times faster in daily use may not improve the experience if it exceeds human reading speed, but if slow thinking makes both sides 10 times slower, a competitor could become “unbearable” while 01.AI remains acceptable. A frontier teacher model can then improve the small model through labels and synthetic data, forming what 李开复 sees as the next mainstream architecture.
  • The “soul-searching” question for model startups has compressed from the past 4-5 years to 2-3 years, and GPUs must undergo explicit return-on-investment accounting like hiring, computers, and travel. Training Yi-Lightning cost $3M, while DeepSeek said its own training cost $6M; if early research is included, 李开复 assumes total investment could reach $10M. With a model lifespan of only about six months, the company must answer how much incremental revenue the new model adds over the old one. “If you burn $400M-$500M a year, even if you’ve raised more than $1B, this question will come very quickly.”
  • 01.AI’s initial commercialization evidence is more than $100M in actual 2024 revenue, overseas To C products roughly breaking even, and multiple software orders worth more than $10M each across domestic gaming, energy, automotive, and finance. 李开复 said it was probably the first of the “four little tigers” founded in 2023 to reach $100M in revenue, while warning that it is “still a long way from an IPO.” His 2025 formulation is to grow from $100M to several hundred million dollars, while reducing model training to once a year, at most twice, so the cost becomes a business expense that revenue can absorb.
  • 01.AI’s customer-acquisition discipline is “don’t fight battles you can’t win”: it should not buy To C growth through sustained marketing, nor pour resources into low-value, hard-to-replicate To B edge projects whose sales costs cannot be recovered. Enterprise orders are worth pursuing if they directly help customers make money, are driven by CEOs who embrace foundation models, possess deep industry know-how, and want to co-build, or are initially unprofitable but allow 60%-80% of the delivery to be reused across follow-on customers. 01.AI is also considering technology-for-equity joint ventures in which industry partners contribute know-how and data, turning a price-squeezing vendor-client relationship into shared upside and downside.
  • 李开复 estimates China’s model capabilities trail the US by roughly six months, yet China has achieved a “90 versus 98” result with less funding and tighter chip constraints; the next contest—business models and AI-first applications—should suit China better. He rejects the idea that startups will all fail, but predicts that “three years from now, no company will be regarded as a foundation-model company,” unless it becomes OpenAI. Winners will shed the technology-era label and become category leaders like Meituan and ByteDance. He acknowledges this is a long-horizon forecast and cannot be answered precisely. The key test for a startup opportunity is whether AI redefines the product—not whether it merely adds a Copilot to Office, search, or an existing platform.

Deep dive

1. The Joint Lab Redraws the Boundary, but 01.AI Was Not Acquired by Alibaba

  • 李开复’s clarification is that 01.AI and Alibaba Cloud have formed an industrial foundation-model joint lab. The people pursuing massive clusters and frontier-scale models are moving into it and becoming Alibaba employees; 01.AI is retaining smaller training, infra, and applications teams and will continue operating as an independent company.

  • What has been explicitly abandoned is frontier-scale-model pretraining, not all pretraining. 01.AI already has Yi-Lightning V2; going forward, it may train once or twice a year, or continue using an existing model. There is only one test: “Is the value it creates worth doing? If not, I’ll use the previous model.”

  • 李开复 compares the new division of labor to a university and a teachers’ college: Alibaba “trains the teachers,” while 01.AI “trains the students,” with “super students taught by exceptionally strong teachers” eventually entering the workforce. The people who wanted to work on Scaling Law and clusters of 10,000 GPUs or more have therefore moved to a platform better able to support that ambition.

  • On a potential acquisition, he said 01.AI “did not seek to be acquired,” while acknowledging that any startup has a responsibility to consider an outcome that best serves investors. A business spin-off is likewise an open option: if it could improve employee motivation without sacrificing control, it could be discussed, but no split has occurred.

2. The Shift From Yi-Large to Yi-Lightning Began Months Before the Rumors

  • In May 2024, 01.AI faced a choice: keep chasing larger models with more GPUs and more data, or build products that could be deployed and monetized. Yi-Large ranked well but was “neither fast nor cheap”; the planned Yi-XL was abandoned in May or June, and the company pivoted to the MoE architecture it had already been exploring internally.

  • Launched in October, Yi-Lightning used a completely different mixture-of-experts architecture. It performed a full tier above Yi-Large, ranking sixth globally at the time, while Yi-Large may have fallen to around 15th-20th; it was several times faster and priced at roughly one-thirtieth of GPT-4o.

  • 李开复 said Yi-Lightning was only slightly behind the latest GPT-4 at the time, while outperforming the GPT-4 available in May. That reset 01.AI’s objective: not to build “the world’s largest and most expensive model with the best performance,” but to achieve the best performance possible while remaining cheap and fast enough.

  • 程曼祺 asked why 01.AI should still build small models when Alibaba’s Qwen already comes in multiple sizes and the two companies are entering a deep partnership. 李开复’s answer: as long as Yi-Lightning cannot be replaced by open source on cost and performance, it deserves to exist. The company is not “finding applications for Yi-Lightning”; it is “finding the best model for applications.”

3. Training-Time Scaling Law Still Works, but No Longer Delivers Startup Returns

  • 李开复’s updated view is not that Scaling Law has stopped working, but that it has entered diminishing returns: “Going from one GPU to 10 might deliver the value of 9.5 GPUs,” while increasing from 100,000 GPUs to 1 million “might deliver only the value of 130,000.”

  • Continuing to burn money on frontier-scale models is suitable for companies firmly committed to AGI, determined to build the world’s largest and best model, and able to absorb the cost. He says that is “absolutely not something a startup can do,” unless it genuinely has an exceptionally long, hard-to-replicate advantage.

  • In China, 李开复 identifies Alibaba and ByteDance as the companies most clearly able to continue in this direction, while leaving room for “perhaps one or two more” entrants later. The logic behind 01.AI’s partnership is straightforward: “If you build a great mobile app, you wouldn’t say you’re going to rebuild Android.” A frontier teacher model will likewise become a foundational operating-system-like capability.

4. Frontier Models Move Into the Classroom; Small Models Become the Students Entering the Market

  • Frontier models still have critical uses: their labels can strengthen post-training, and they can generate synthetic data better suited to student models. 李开复 stresses that synthetic data is not more valuable merely because it replaces real data. Its real value is generating a higher-quality 20T tokens for a model whose original 20T-token corpus has saturated, then pushing the model’s capabilities up another level.

  • He expects the mainstream market to consist of capable-enough models such as Yi-Lightning, smaller Qwen variants, DeepSeek, and GPT-4o mini. They do not need to be state of the art; they need to be continuously taught by stronger teachers, then carry mass adoption through speed, cost, and user experience.

  • 李开复 relayed an explanation he heard from friends for Opus and Sonnet: Opus is too large and too slow; if sold openly, buyers might distill it into competing models. It is better kept as a teacher for training Sonnet, which is the model sold commercially. This was hearsay, not a conclusion he had personally verified.

  • He also mentioned that the names GPT-4.5 and GPT-5.0 were still undecided, although the models had already been built. Internal testing showed improvements, but perhaps not enough to justify the associated latency and cost. Whether or not it is sold, it will serve as a teacher that “raises all the GPT small models another level.”

5. Test-Time Scaling Law Turns Inference Speed From User Experience Into a Cost Lever

  • 程曼祺 noted that Scaling Law could also shift toward inference time, or test time. 李开复 sees this as a separate curve and believes “slow thinking, long thinking” may drive the next breakthrough, making fast inference more valuable rather than less.

  • If 01.AI is normally five times faster than GPT-4o but both models already produce tokens near or above human reading speed, users may not receive five times the benefit. If slow thinking makes both sides 10 times slower, the competitor could become “unbearable for many applications,” while 01.AI remains within an acceptable range.

  • A fast inference engine also allows more experimentation and accelerates the search for better slow-thinking methods. 01.AI’s choice to be “small, fast, and cheap” is therefore not only defensive cost control, but also an early bet on test-time Scaling Law.

6. The Industry’s Soul-Searching Question Moves From 4-5 Years to 2-3

  • 李开复 summarizes the AI 1.0 path as papers, competitions, orders, and expansion, with financial statements examined only at the end. In AI 2.0, everything is accelerating: it took only a year to move from faith to doubts about Scaling Law, and capital will ask earlier whether technology can become commercial value.

  • Even companies from the previous generation, such as SenseTime and Megvii, took roughly 4-5 years to face financial scrutiny after excluding their quieter early periods. This cycle may allow only 2-3 years. If a company burns $400M-$500M a year, “even if it has raised more than $1B, this question will come very quickly.”

  • The sequence of questions is: Does the company understand how to operate commercially? How much revenue does it have? Can it grow? Can it control costs? It must then narrow losses and move from single-point profitability to profitability across multiple lines. Otherwise, even with revenue growth, losses at 5x, 10x, or 20x revenue will be difficult to accept.

  • GPUs therefore have to be treated as a business expense. Training Yi-Lightning cost $3M, while DeepSeek said its own training cost $6M; if early research brings the assumed total to $10M and a typical model lasts about six months, the company must calculate what the upgrade adds over the old model, just as a household with an existing car decides whether to replace it immediately.

7. The First Commercial Rule Is to Reject Battles You Cannot Win or Make Self-Sustaining

  • 李开复’s first discipline is “don’t fight battles you can’t win.” Companies should not invest merely because a field looks technologically advanced if it lacks PMF, requires years of market education, or is already dominated by a giant with an overwhelming lead.

  • The danger in To C is buying users through advertising, only to see growth disappear when spending stops. Even if free users remain active, they consume GPUs. With domestic To C monetization difficult and giants entrenched, maintaining rankings before organic growth or a path to monetization emerges only demands continued cash injections.

  • In To B, edge-case features create little value for customers, generate little willingness to pay, and are difficult for suppliers to deliver well; the vendor may ultimately fail to recover even its sales costs. The projects worth pursuing are core initiatives that directly help customers make money, or strong partnerships driven by visionary CEOs with deep industry know-how who understand where foundation models fit and are willing to embrace them.

  • Another viable order is initially unprofitable but produces reusable delivery assets: 5, 10, 20, or even 100 subsequent customers can reuse 60%, 70%, or 80% of the first delivery. In 李开复’s view, reusability determines whether project revenue can ultimately become product profit.

8. $100M in Revenue Is Only the Starting Validation; 2025 Targets Several Hundred Million and Explainable Costs

  • Founded in 2023, 01.AI was probably the first of that year’s “four little tigers” to generate more than $100M in actual revenue in 2024, 李开复 said. He tempered that with: “$100M in revenue doesn’t mean much; we’re still a long way from an IPO.” But reaching that level in the first operating year shows that commercial discipline is beginning to take hold.

  • Overseas To C products have roughly broken even and may become profitable. In China, 01.AI has won business in gaming, energy, automotive, and finance, mostly software orders worth more than $10M. 李开复 uses order size as evidence of value: being able to close eight-figure software contracts shows that the company is creating real value for customers.

  • For 2025, he said he is confident of “several-fold growth,” taking revenue from $100M to several hundred million dollars. Alongside expanding existing verticals, 01.AI will co-build with industry partners in selected areas. One model is a joint venture in which the partner contributes know-how and data while 01.AI contributes technology as equity, with both sides sharing the outcome.

  • The desired financial structure is ultimately straightforward: revenue grows several-fold, the model is trained once a year, at most twice, and an external teacher model keeps it from falling behind—or even keeps it in the first tier. If GPU training costs can be reasonably allocated against revenue, the financial statements will be legible even to investors outside the foundation-model industry.

9. The Adjustment Followed Months of Deliberation, and Re-exposed the Boundaries Around Talent and the CEO

  • 李开复’s timeline begins in May 2024, when the idea emerged; the path became clearer in the third quarter, after which 01.AI and Alibaba agreed on the joint-lab structure and began implementation over the past month. The change came from updating industry conditions, understanding, and execution—not from a sudden, rumor-triggered retreat caused by funding pressure.

  • On departures among mid- and senior-level staff, he said the core founding team remains largely intact, but every company in the sector is recruiting with “sky-high” offers from Big Tech. Some people cannot resist the financial incentives; others genuinely want to train AGI-scale frontier models. Since 01.AI is no longer offering that environment, it cannot force the latter group to stay.

  • 程曼祺 asked whether failing after achieving fame and success would damage a founder’s résumé. 李开复 said he carries no such baggage: “The AI era I had waited more than 40 years for” had finally arrived. Not doing what he is good at and taking a shot would be the lifelong regret; a CEO who regrets making adjustments “is someone who is not qualified to be a CEO.”

  • His review of entrepreneurial ability is that leaders must hold firm to what matters, respond quickly to changes that have already occurred, and adjust boldly based on their view of the future. The original dream of becoming “Microsoft in the AGI era” has not been declared dead; the route has simply changed from models first to applications first. Anyone can look to the stars, but it matters more to keep one’s feet on the ground.

10. The “Foundation-Model Company” Label Will Disappear in 3 Years; AI-First Will Determine Who Breaks Out

  • 李开复 does not believe China’s foundation-model startups will all fail. Each has money and smart teams and will find a direction. But he maintains that “three years from now, no company will be regarded as a foundation-model company,” unless it truly becomes OpenAI. The other winners will become leaders in overseas To C, domestic To B, or specific industries. He also acknowledges that this is a long-horizon forecast and cannot be answered precisely.

  • If only existing technology giants ultimately benefit, he says that would mean AI-first applications failed to produce the expected disruption. If every change is merely an upgraded version of Douyin, Baidu, or Taobao, Big Tech will extend its success; if product categories are redefined, incumbents’ legacy businesses will leave an opening for new companies.

  • He uses the transition from PCs to the mobile internet to explain the dividing line. Search changed relatively little, but transport, video, payments, and local services were rebuilt around location and constant portability. If phones had remained limited to H5 versions of old websites, a new generation of mobile giants would never have emerged.

  • The test for AI-first is whether the application could exist at all without a foundation model. Office plus Copilot does not qualify. A structural rewrite means moving from “humans write and AI helps” to “AI writes and humans adjust,” from keywords and website lists to questions and answers, and eventually to social circles that include both people and AI.

11. China Enters the Application Race at “90,” Its Stronger Field; 2025 Will Still Be Full of Non-Consensus Opportunities

  • 李开复 believes China’s model capabilities trail the US by roughly six months, but China has built cheaper, smaller models under chip restrictions, lower valuations, and less funding—good enough for a “90,” versus roughly a “98” for the US. China was disadvantaged in the first model-technology race; the next race, in business models and applications, better matches its strengths in execution and finding PMF.

  • His most certain predictions for 2025 are, first, that both China and the US will produce large numbers of AI applications, many from Chinese startups and Big Tech; and second, that there will be more surprises, because “many things once regarded as truths will soon be overturned.” Forecasts must leave room for unknown breakthroughs.

  • Non-consensus To B opportunities will not be broad categories such as finance or insurance, but granular verticals. They must depend on foundation models, do more than save a little money, and help companies make money—possibly even doubling a category leader’s revenue or employee productivity. This kind of fine-grained PMF is where 李开复 sees the next phase’s most valuable opportunities.

  • To peers, 李开复 cited 王慧文: “Every one of us is a warrior; we should encourage one another.” For super agents, he reserves the human role for unprecedented ideas, combinations that require synthesis, predictions that cannot be derived from data alone, human warmth, and new opportunities created with AI. If AI frees up time, he will continue doing the things he loves that AI cannot replace—and spend more time with the people he loves.