
朱亦博
Frontier Insights
Core Thesis: AI Infra is no longer support plumbing—it is the strategic battleground for LLM margins. At a 10k-GPU scale, every 10% gain in utilization yields ~¥10M monthly, making in-house infra non-negotiable once scale hits critical mass.
Strategic Decisions: Pure intermediate tooling faces commoditization and price wars. StepFun prioritizes deep hardware-model-system co-design, actively anchoring its Moat around vision-reasoning workloads and domestic silicon integration.
Risks & Warnings: Standalone infra startups lacking proprietary models or hardware control will be crushed. Long-term survival demands vertical efficiency, not horizontal abstraction.
Key Views & Dialogues
Everything About AI Infra | A Conversation with StepFun Co-founder 朱亦博
- 🗓️ Date:
2025-08-02| 🎙️ Show:42章经
Large models have moved AI Infra to the center of model competition: renting 10,000 GPUs costs about $100M monthly, so a 10% utilization gain can save or generate roughly $10M, making in-house teams more compelling at scale. Independent vendors stuck between commodity hardware and models risk price competition, while StepFun is pursuing stickiness through visual reasoning, Chinese-chip adaptation and deep co-design, with the next paradigm and model-chip integration still unresolved.
View Dialogue Notes & Key Takeaways
Large models have pushed AI Infra from a back-office cost-cutting tool to the center of model competition, creating what 朱亦博 sees as an industry window that opens only once every 10 or 20 years. Search engines turned Google into an Infra company because of their massive data and compute demands, and large models are repeating that process; the underlying protagonist has simply shifted from CPUs to GPUs, with compute, networking, and storage all custom-built around the model.
The economics of AI Infra are highly quantifiable, and the larger the scale, the more compelling it becomes to build an in-house team. 朱亦博’s math: renting 10,000 relatively expensive GPUs costs about $100M a month, so a 10% utilization improvement can save or generate roughly $10M monthly; at smaller scale, a general-purpose cloud baseline may be sufficient, and even 1M DAUs may not justify building a full Infra stack.
An independent AI Infra vendor positioned only between commodity hardware and commodity models will struggle to build a durable moat. A single optimization may lead for a few months, but “there is no technology that cannot be caught within a few months,” making a slide into price competition likely; the real way out is to move toward hardware or models and build stickiness through proprietary compute, first-party models, or deep co-design, because “you should not be the person stuck in the middle.”
In 朱亦博’s view, o1 and reinforcement learning shifted the industry’s primary metric from training MFU to decoding speed and cost, though consensus has yet to fully form. DeepSeek initially optimized for low inference cost, leaving its training MFU relatively low, and its base model was not necessarily ahead in the first half of 2024; after o1 introduced test-time scaling in September 2024, low-cost inference translated directly into reinforcement learning running several times faster, a key condition for it to produce R1 first. “The most important thing is always the choice of direction.”
Model competition is not a single event for algorithm teams, but a “three-legged stool” of algorithms, systems, and data. If two teams train with 5,000 GPUs for 3 months and Infra raises efficiency by 20%, the model can learn 20% more data under otherwise identical assumptions, potentially improving the final result; 朱亦博 goes further, arguing that model architecture determines operating cost, systems teams should be deeply involved in design, data teams should own outcomes, and algorithm teams should focus on training methods.
StepFun is pursuing two-way vertical integration between a visual reasoning model and Chinese-made chips. 朱亦博 says its new model should be China’s first visual reasoning model in the hundreds-of-billions-parameter range available for commercial use by third parties, capable of analyzing images directly to decompose tasks rather than converting them to text first; StepFun will offer free commercial licenses, share weights, and help adapt the model to all Chinese chipmakers, while architectural innovations push inference costs on domestic cards down to levels competitive with Nvidia solutions.
The next paradigm shift may come from unifying multimodal understanding and generation, while specialized capabilities and Agent applications still have a window of opportunity. Using roughly 2 years as a rule of thumb, 朱亦博 places the next shift after GPT-3.5 in 2022 and o1 in September 2024 perhaps in 2026; Claude has differentiated itself through investment in code data, while Agent companies and model vendors are in a relationship where “there is co-existence, but they are also hurting each other”(有共生,但是又在互相杀伤). His long-term constant is The Bitter Lesson: “In the long run, the winner is always the method that can make the most use of computation.”
🔗 Original source & video: Everything About AI Infra | A Conversation with StepFun Co-founder 朱亦博