99: MiniMax Founder Yan Junjie: Never Apply the Mobile Internet Playbook to Foundation Models
99: MiniMax Founder Yan Junjie: Never Apply the Mobile Internet Playbook to Foundation Models
Summary
- MiniMax has shifted its core objective from revenue, growth and user scale to technological iteration. Yan Junjie argues that better models can create better applications, but a larger application footprint will not directly improve model capability in return; ChatGPT’s DAU may be 50–100x Claude’s, yet its model is nowhere near 50–100x better, and too many users can even slow R&D and iteration.
- MiniMax-01 was open-sourced for the first time primarily to shorten the R&D feedback loop and begin building a technology brand. Public models expose both strengths and weaknesses faster, while helping attract enough high-caliber talent. Looking back, Yan says that if he were starting over, open source should begin on “day one”; he also sees no reason to hide a stronger version indefinitely, given how short a model’s shelf life is.
- MiniMax-01’s core architecture replaces Transformer’s softmax attention with linear attention to support much longer contexts. This provides the computational foundation for long-term memory in a single Agent and communication among multiple Agents, but capabilities such as tool use and planning have not yet been optimized; Yan also made clear that the next version should not be assumed to have reached o3.
- “Technology-driven” means prioritizing the next capability step when resources conflict, rather than fixing every local problem first. When Hailuo Video faced rough interfaces and missing features, the team still chose to “listen to the algorithm”; Yan attributes Hailuo Text’s weaker performance to a period when the company failed to stick to that principle. Product, operations and algorithms should all serve the continued improvement of technical capability.
- AI products cannot simply import mobile internet logic around user growth, paid acquisition and advertising. Yan believes products that depend heavily on promotion are probably not sufficiently compelling; products should first prove their own value, with mature businesses measured through retention and LTV, while overseas AI products are currently easier to evaluate through paying users and subscription conversion.
- MiniMax missed its 2024 product and revenue targets not because the targets were too ambitious, but because they were built on mobile internet growth curves. In 2025, the company first assesses what R&D can achieve and what product changes that will create, then plans business, revenue and budgets, rather than setting business targets first and reverse-engineering the technology.
- The organizational overhaul follows the same trade-offs: everyone should stay hands-on, propose solutions and make decisions based on current facts rather than experience from former employers. Yan acknowledges that personnel changes were delayed and the company’s goals at one point oscillated between revenue and growth; he attributes his own development to letting go of ego, admitting mistakes and making timely trade-offs. The biggest organizational challenge now is continuing to attract strong people, while it remains uncertain whether MiniMax can reach professional-level standards in some specialist fields in 2025.
Deep dive
1. Product awareness and technology branding remain weak spots
- After MiniMax-01 launched, overseas users first asked when the company would update its voice, video and music models. That suggests MiniMax is better known overseas for other modalities, while also exposing the fact that its text models had not previously established a technology brand abroad.
- Domestic feedback fell into two camps: people in algorithms and academia saw, perhaps for the first time, that a non-standard Transformer architecture could work at meaningful scale; partners and friends saw MiniMax beginning to recognize the importance of a technology brand.
- Yan believes technology branding is an area the company underinvested in over the past 2–3 years, and a key reason for this open-source release.
2. Open source is first about shortening the R&D feedback loop
- Yan expects technological progress to remain rapid for at least the next 2 years, as it has over the past 2 years. What matters is not only how good the current model is, but how quickly it is improving.
- Open source encourages strengths and exposes weaknesses faster; the feedback loop may even be more efficient than putting a closed-source model into a product and waiting for user feedback.
- The other reason is technology branding. Compute, data and capital are hard constraints; enough high-quality people are a critical soft constraint. Open source is MiniMax’s starting point for building a technology brand.
- Asked why the company did not open-source earlier, Yan said that, given the choice again, opening up sooner would have been more rational—perhaps even from “day one.”
3. No stronger closed-source version was deliberately held back
- MiniMax already uses MiniMax-01 internally and is not hiding a separate, stronger version. It will continue modifying the open-source version, but sees no need to preserve a better model solely for itself.
- Yan’s reasoning is that models have very short shelf lives: every model could be obsolete a year later, so there is little long-term value in hiding the stronger model available today.
- He also believes open source can reduce the security and operational burdens that come with closed-source models in customer engagements, making it easier to meet customer needs. He went further: “If I were OpenAI, I should open-source too.”
- In his view, OpenAI’s core advantage is closer to the ChatGPT brand and user mindshare than to GPT-4o being inherently better than Claude 3.5; when models are broadly comparable, one released by OpenAI is also more likely to become the trend.
4. DeepSeek’s lesson was branding and focus
- By the time MiniMax recognized the importance of a technology brand, DeepSeek had not yet open-sourced V3, so MiniMax’s move was not a direct response to DeepSeek V3.
- Yan has known 梁文锋 for a long time and considers 幻方 one of the “one or two” quant firms with the strongest reputations. 梁文锋’s understanding of brand value had a major influence on him.
- DeepSeek had no product in its early days and could focus its resources on models. Yan also notes that DeepSeek later launched an App, so it is inaccurate to say simply that it only builds models and not products.
5. User scale is not the same as model intelligence
- Yan uses ChatGPT and Claude as examples: ChatGPT’s DAU may be 50–100x Claude’s, but its model capability is clearly not 50–100x higher; the two models are broadly comparable.
- Alibaba’s Tongyi and Doubao in China make the same point: user numbers have no simple relationship with model intelligence. The industry’s past misconception was that more users and more feedback would accelerate model improvement, leading companies to spend heavily on traffic acquisition.
- Capability gains do create new products such as AI Coding and video generation, but these capabilities were not developed by first gathering massive user feedback and iterating from there. Many directions begin with a defined target and benchmark, followed by training against that target.
- More users can create product value, but they do not automatically make a model smarter; too many users may even constrain R&D and iteration speed.
6. Recommendation-system methods cannot be transferred directly to AGI
- In recommendation systems, it is often impossible to determine in advance what content is right or wrong, so filtering through large volumes of content, experiments and A/B metrics is relatively effective.
- Transferred to foundation models, that approach becomes dividing scenarios into categories, adding data, fixing cases and running A/B tests one category at a time. Yan believes this can patch local experiences, but it is not a way to raise AGI’s ceiling.
- His preferred approach is to define capability levels, training data, inference processes and reasoning complexity first, then use technical methods to approach the target.
- He has seen companies launch good models, gain users and then begin iterating around user needs—only for R&D speed to slow and a newer model to overtake them.
7. Model and product are not a two-way flywheel
- Yan realized in March or April 2024 that technology and product should not be conflated. Technology’s job is to keep raising the capability ceiling, potentially changing existing products or creating new ones.
- Product’s job is to satisfy user needs and generate commercial behavior; once a product exists, it does not naturally make the model better.
- His summary is simple: better models can produce better applications, but better applications—or applications at greater scale—do not directly produce better models.
8. Being technology-driven means prioritizing the next capability step
- When a product launch generates a long list of unsatisfactory cases, a team can either fix every local problem or tolerate some of them temporarily and focus on making a larger capability leap.
- MiniMax chose the latter. Yan cites Hailuo Video: by traffic, it may already be the world’s largest video-generation product, yet its interface was rough and it initially did not even have an English-language page. Overseas users asked for Runway-like features and Kling’s App, but the team still chose to “listen to the algorithm” (“听算法的”).
- His view is that once the technology is no longer leading, the product may have no future. Product, operations and algorithms should therefore all help improve algorithmic capability.
- Entertainment products such as Xingye also require sophisticated algorithms. In 2023 the team sometimes agonized over the trade-offs; by 2024 it hesitated much less.
9. Hailuo Text shows what happens when technology-driven principles slip
- Yan believes Hailuo Text failed to take off mainly because the team did not remain technology-driven. When users raised problems, he considered patching corner cases instead of looking for the underlying path to a real capability improvement.
- He says Hailuo Text was behind competitors in March but no longer was by May. 曼祺 argued that Hailuo had already lost category mindshare by then; Yan said “no,” while acknowledging that he had already judged Doubao the likely winner because its experience was clearly better.
- MiniMax ultimately did not spend aggressively on user acquisition like other companies because user numbers have no particularly strong connection to technical capability gains; the product should be treated as an independent business.
- On Hailuo repeatedly pushing notifications to existing users, he acknowledged that this was a consequence of not holding firmly enough to technology-driven principles, and partly attributed the wavering to concerns about team members’ feelings.
10. The “chasing hot trends” narrative ignores the timeline
- Some describe MiniMax as moving from virtual humans toward Character.AI, restarting Hailuo after Kimi took off, doubling down on video after Sora appeared, and then following DeepSeek into open source. Yan believes this account fundamentally misunderstands the company.
- Glow and Xingye emerged around the same time as Character.AI, with the mobile versions arriving earlier. Hailuo launched 2 years ago but failed to gain traction for more than a year; only after Kimi performed better did outsiders assume Hailuo had been restarted.
- The video project began with the idea of bringing Talkie characters to life. After Sora appeared, the team realized the opportunity was larger than originally imagined and made the product more general-purpose.
- His summary of the long-term path is that foundation models first create leisure and entertainment applications, then gradually enable higher-productivity applications; open source is tied to technology branding and the R&D feedback loop.
11. AGI is a directional long march
- Yan believes no one currently has the ability to define AGI precisely. MiniMax writes “Intelligence with everyone” on its walls; what the company can say with confidence is that intelligence matters and can continue improving.
- He compares the effort to the Long March: the final destination may be unknown, but higher intelligence is probably valuable.
- The technology, products, data and capabilities accumulated at each stage should help the company enter the next. A startup has fewer resources than overseas companies and China’s major internet firms, so it cannot lay out the entire road in advance; it can only advance one step at a time.
12. Hitting a scaling-law wall is not a reason to quit
- Yan does not fully agree with the claim that the industry went from believing in scaling law to doubting it within a year. Making a company or technology 10x better is inherently difficult and does not happen automatically.
- ByteDance can pursue multiple directions simultaneously and use internal competition to select winners; startups have fewer chances to make mistakes. But limited resources can also force startups to innovate in algorithms, organization, business or direction.
- His view is that companies should not abandon the effort when scaling law hits a wall, but look for ways to keep technology improving. Whether the method is still called scaling law is secondary; the key is to sustain progress.
- He does acknowledge a failure condition: if no method can ultimately be found, the company should indeed shut down. But while there is still a chance, it should keep searching.
13. “Belief” has to show up in resource allocation
- 曼祺 questioned whether “faith” is something that can be realized within a year. The discussion ultimately centered on whether a company’s long-term direction is consistent with its short-term spending.
- Citing OpenAI and Anthropic, 曼祺 said they spend mainly on compute and talent rather than large-scale marketing. Yan agreed that resource allocation should reflect technology-first priorities.
- He prefers the word “belief”: first, believing that technology can keep improving; second, when resources are limited, choosing the things that genuinely advance technology while keeping the company moving forward.
14. The standard for an Agent is handling complex tasks
- Yan sees Agents through two lenses. Technically, as AI becomes stronger, it should handle more complex tasks; socially, AI should not merely reply instantly but work over a longer period and deliver a complete result.
- A multi-step task can be completed through a single o1-style reasoning pass, broken into a predefined workflow, or handled through collaboration among multiple Agents.
- He compares this to a task assigned by his mother: an instant reply is one format, while spending 2 days preparing and then delivering the result is closer to complex-task execution.
- His standard for a complex task is professional-level performance in a professional field. MiniMax is currently focused mainly on the digital world, not because it believes Agents should never enter the physical world, but because it does not yet have the ability to do so.
15. MiniMax-01 has completed only the architectural prerequisite for Agents
- Agents need very long memories and contexts, while multiple Agents need to exchange large amounts of information. Yan believes MiniMax-01 already has the architectural conditions to handle long contexts, and says it “should also be the only architecture in the world currently able to do this,” retaining the qualifier “should.”
- At the capability level, substantial work remains. Tool use and planning, for example, have not been optimized particularly well.
- So “opening the Agent era” mainly means that the computational architecture is in place; it does not mean the current model can already perform professional-grade Agent work.
- The next step is to build benchmarks around these capabilities, improve scores while preserving generalization, and follow a two-step roadmap: “underlying computational architecture” first, then “all the capabilities being good enough.”
16. Linear attention is a major variant of Transformer
- MiniMax-01 can still be understood as a major variant of Transformer, but it replaces softmax attention with linear attention.
- Google’s advantage is that its TPU, training framework and algorithms can be co-designed. MiniMax cannot customize TPUs and must modify algorithms and software on standard hardware, making implementation more complex.
- Google’s implementation is closed source, so Yan can only speculate that it may use sliding-window attention. MiniMax does not use a sliding window; it uses an approximation algorithm that still incorporates the content into computation to improve efficiency.
- He does not see linear attention as an industry-wide consensus, comparing it with DeepSeek V2’s MLA. Different routes may all work; the key is to choose one that fits the company’s conditions and pursue it consistently.
17. Long-context processing is about time and memory
- Yan believes “long context” is more accurate than “long text.” Humans experience the passage of time; AI must solve how to process memories that keep growing.
- A single Agent should not be viewed merely as one conversation, but as a process that continuously receives inputs and generates outputs along a timeline—perhaps even an “infinite conversation.”
- Multiple Agents require not only strong individual Agents but also efficient communication between them. Yan mentions MCP and notes that the communication content itself may be very long.
- The foundational capabilities for multi-Agent systems are therefore individual capability, coordination capability and efficient long-context processing.
18. Benchmarks are a roadmap for capability
- Yan believes many of the targets that have driven major advances in AI capabilities were defined by academia through benchmarks.
- He cites SWE-bench, recalling that the score was around 10 points a year ago and is now above 70. The improvement in coding capability came from clear targets and training methods, not simply from more users.
- For its next-generation models, MiniMax will strengthen Coding and planning while continuing to tackle complex benchmarks for multimodal Agents.
- Architecture determines the computational pattern; capability is the concrete set of parameters learned according to that pattern. Assessing an architecture’s ceiling requires both judgment and extensive experimentation and collaboration.
- One measure of improving R&D capability, he says, is whether the team designs better experiments than it did several months ago when the search space remains very large.
19. No rush to distill an o1 label
- Yan believes it is not difficult to build something that “looks like o1”; distilling several thousand sets of o1 data might be enough. MiniMax has run related experiments internally, and many papers on the subject have appeared recently.
- But the company’s current business does not depend on having an o1 model, so there is no need to distill one and then publish an article or press release.
- He cares more about making the training, capability gains and generalization process robust. At least for now, MiniMax does not need to claim it has built o1.
- He separates architecture from capability: architecture provides the computational conditions, while the reasoning capability represented by o1 belongs to the capability benchmark; improving that benchmark may also require changes in training methods.
- Asked whether the next version would reach o3, he explicitly said no, only promising stronger Coding and planning, and stressing that capability must be judged against specific benchmarks.
20. One month early or late is not the point
- No company can be first in every direction because capability and resources are limited. Over the past 2 years, the companies that ultimately performed best were often not the earliest starters; they were the ones that executed most rigorously and made the fullest use of their strengths.
- Yan therefore does not see being 1 month early or 1 month late as the core issue. Skipping foundational work can produce results faster, but may prevent the company from doing its best work.
- Inference scaling did not begin with o1. Best-of-n, repeated sampling, batching and tree structures all existed earlier; o1’s main advance was turning multi-step search into an end-to-end model that could be optimized globally.
- When 曼祺 asked whether memory and multimodality should be moved forward, Yan stressed that the company should not skip foundational work simply because later-stage projects can show results faster.
21. Distillation is one path; whether it is a shortcut is unclear
- Yan believes Chinese companies moved faster on o1 partly because there are more companies and a broader belief that distillation is a viable path. If Google or Anthropic chose distillation, they might also move quickly, but he assumes they have their own judgments and roadmaps.
- Asked whether distillation is a shortcut, he rejected the certainty of that framing. Distillation is definitely one path; whether it qualifies as a shortcut is a matter of perspective.
- On the potential cost, he cited the “alignment tax” in text models: aligning a model to GPT-4’s outputs can constrain its capabilities. He did not go further and claim that distilling o1 would necessarily create the same cost.
22. Multimodality is necessary for professional-grade Agents
- If an Agent is to do a doctor’s job, it needs to see; if it is to solve geometry problems, it needs to understand shapes. Multimodality can be native to the model or added afterward, but it cannot be avoided.
- Yan did not give an absolute answer to the view that video generation contributes less to intelligence than continued iteration on the language base model. Different people understand intelligence differently; in his view, generating content is itself part of intelligence.
- He uses Cursor to show that a Coding Agent does not need an o-series model: Cursor’s early success was built mainly on Claude 3.5, and integrating o1 into GitHub did not make the experience clearly better than Cursor.
- Hailuo Video’s weekday usage is far higher than its weekend usage, and some users are willing to pay, indicating real productivity value. But image quality, physical consistency and generation speed still have major issues; it has only just reached practical usability.
23. Information retrieval may be the next major Agent use case after Coding
- Yan believes Coding is one of the earliest Agent applications to reach real-world use, while information retrieval may be another category to emerge quickly.
- Users may want to see the 10 best papers in a field every day, or know whether Jay Chou is holding a concert in China. An Agent can continuously track such explicit needs instead of waiting for users to open a content platform.
- On the comparison to a “new Toutiao,” he warned against applying mobile internet product methodology to AI. AI may handle both distribution and supply, while model capabilities and the way supply is obtained continue to change.
24. Heavy reliance on promotion may signal a product that is not yet fully formed
- Yan says Hailuo Video spent nothing on promotion overseas and did not promote in China. His real point is that if a product depends heavily on promotion, it probably is not quite right yet.
- He sees the promotion histories of Glow, Hailuo, Xingye and Talkie as evidence of a changing understanding: early on, the company did not know how to promote; later it learned the methods of major internet firms; eventually it recognized that those methods do not apply to every company.
- For its own products, retention and LTV can help determine how to operate the business. But if a startup can only rely on mature metrics such as these, that may indicate insufficient innovation.
- AI products have different growth logic from mobile internet products because model capability, product form and the way supply is created are all changing.
25. Alignment determines a model’s character
- In the Copilot-versus-Agent discussion, Yan says the two products do use different approaches. ChatGPT is more like an assistant helping users complete tasks, while Claude sometimes feels more like a friend with emotional intelligence.
- As an example, ask a model to choose a number between 1 and 100, then tell it that you will not speak to it for that many days. Claude might ask for another chance and choose a smaller number, while ChatGPT might not understand the interpersonal implication.
- He links Claude’s character to Anthropic’s system spanning values, constitutional principles and alignment data. The model may not have been directly trained to behave this way, but the overall constitution and alignment methodology can produce that pattern.
- He believes most Chinese models are still primarily aligned to OpenAI’s outputs, while Doubao is beginning to develop its own character—for example, offering users 3 poems to choose from when asked to write poetry. MiniMax is also trying to define the traits its models should have first, then apply different alignment methods to different products.
26. Both model companies and application companies can work
- Yan does not believe there are companies in the real world that only build models and never applications; Claude also has some applications. But companies that only build applications and not models certainly exist, including Polaris and Kessler.
- MiniMax’s pattern is usually to build a model first, with products then emerging from technological changes. Different technology stages may correspond to different products, but product scale does not feed back into model improvement.
- He divides companies into those that build products on existing technology and those that build products on future technology. Both approaches are meaningful; MiniMax is better suited to the latter.
27. In 2025, the focus shifted from business targets to technical and organizational iteration
- MiniMax fell short of its product and revenue targets for 2024. Yan says the problem was not that the targets were too high, but that they were still based on the growth curves of the early mobile internet.
- In 2025, the logic changed: first assess what R&D can achieve and what product changes that will create, then set business plans, budgets and revenue targets, rather than defining the business target first and reverse-engineering the technology.
- Organizationally, the basic standard is not to wait for reports or make decisions for others, but for everyone to propose solutions, understand the details and make judgments based on current facts rather than copying experience from ByteDance, Pinduoduo or previous employers.
- Yan acknowledges that personnel changes were delayed and the company’s goals at one point oscillated between revenue and growth. The company is now more unified around putting technological iteration first.
- He summarizes his own change as letting go of ego, admitting mistakes, making timely corrections and having the courage to make trade-offs. The biggest organizational challenge is continuing to attract strong people; he thinks MiniMax may reach professional-level standards in some specialist areas in 2025, but is not certain.