Do Models Actually Eat Applications?—From Canva and Figma to Meitu: The Two Fates of AI Application Companies
Do Models Actually Eat Applications?—From Canva and Figma to Meitu: The Two Fates of AI Application Companies
Summary
- Zhuang Minghao’s core framework for this episode is a one-line verdict: “Models provide capabilities, applications organize outcomes, and markets price the future.” His conclusion rejects the binary take: models eat isolated features such as background removal, outpainting and background replacement—capabilities that already seem native to today’s multimodal models—but “they cannot eat result delivery in complex scenarios, not remotely.” The companies that survive will be those that organize models, use cases, workflows and payments into “outcomes users are willing to buy repeatedly.”
- The clearest debuff case in the multimodal battlefield is concrete: Figma and Canva have both been repriced by the market. Figma rose from $10-20B after listing to a peak that may have approached $70B, but is now back at roughly $15B—down more than 88% from the peak and probably still below its IPO offer price—even as revenue growth should have exceeded 40% for 3 consecutive quarters and may be accelerating quarter by quarter, from roughly 42% and 43% to 45% and 46%. Canva’s secondary-market valuation fell from roughly $50B at the peak to about $32B, while its 2026 revenue-growth forecast was cut from 30% to 20%.
- The same battlefield is also packed with positive buffs: even as Gen-4.2 and Veo 3 appear to “obliterate everything,” the top 10 video-model companies have raised capital at a furious pace for more than 6 months, and Aishi Technology has begun applying for a Hong Kong listing, meaning a Hong Kong-listed video-model company may soon appear. PixVerse, Hailuo, Sora and Vidu are positive cases, as are HeyGen, Lifelike, OiiOii and Higgsfield in video Agents; Meshy and Tripo raised large rounds in the first half, and he expects them to consider IPOs soon. The coexistence of debuffs and buffs shows how “unusually complex and alive” the multimodal ecosystem remains—if anything, he thinks the US may be more interesting.
- The market’s 4 fears can be broken down systematically: models absorb single-point features; costs come before revenue; experimentation is not retention; and growth is not realization. AI applications have the lowest gross margins, with token, compute, operating, brand, model-call and product-rebuild costs arriving first. That also explains why traditional active-user and retention metrics “really are not that suitable for the state of AI applications at this point,” and why everyone is telling a “bigger story.”
- The number in Meitu’s earnings that surprised him most was that its in-house model generated 96% of the final delivered content, consistent with third-party API and token-call costs accounting for only a single-digit percentage of revenue. In the first half of 2025, revenue was RMB2.21B (+22.1%) and net profit RMB635M (up nearly 40% YoY); productivity apps reached 33M MAUs (+43.5%) and 2.35M paid subscribers (+29.8%), with a paid conversion rate of about 7.1% and related ARR of RMB620M—“on the verge of joining the $100M ARR club.” The legacy base and the second curve are running in parallel.
- On product strategy, he offers 2 actionable tests: letting users choose a model is “a detour,” an expedient compromise for the moment; and the XVC Liu Jing team’s test, in which nearly 60% chose “New Lab,” looks like a joke but is rational. Once open-source models reach the required capability, user memory and the Harness Loop become more important, and with “Own your intelligence,” applications and models are no longer a simple binary.
- His final advice to mature application companies is “both/and,” paired with a counterintuitive rule: the opening line of 郭星’s essay “Meitu Survives Like a Startup” is, “Mature legacy products cannot funnel traffic to new products.” The open questions remain: whether the productivity second curve—roughly 18% of revenue—can become the main curve; the unit economics of token consumption, since growth in token usage does not equal profit growth; resource fragmentation under the incubation structure; and long-term retention and payment in high-ARPU overseas markets.
Deep dive
1. “New Lab” Is Not a Joke; the Binary No Longer Works
- He opened by revisiting the school-bus meme that keeps resurfacing: in the mobile-internet era, the bus represented apps for cameras, rulers, weather and photo albums, while the train running them over was an iOS or Android system update. In the AI era, the bus has become “the application companies built on top of large language models,” and the train is naturally the model companies. He first used the image during the GPT-4 era—3 years ago, several eras ago—and it has been debated through every subsequent cycle of hype and consensus.
- A question in XVC’s Liu Jing team’s “2026 National AI Application Midterm Exam” asked how to define a company that simultaneously builds its own model, trains models for other companies and offers a To C product. Nearly 60% chose D: “New Lab.” Zhuang Minghao’s view is that the answer looks like a joke but is rational. As open-source models reach a certain capability threshold, the latest and most expensive models cost more than expected, and user context memory and the Harness Loop become increasingly important, “Own your intelligence,” as the Sequoia conference put it, means that proprietary, open-source and modified-and-trained models all become part of one unified problem.
2. The Multimodal Battlefield Carries Debuffs and Buffs at Once
- There is still no consensus on how to divide the battlefield. One camp, represented by 梁文锋, believes AGI can be achieved through language alone and that multimodality is an unimportant peripheral layer. In China, however, the ecosystem spanning video models, video Agents, world models, 3D, voice and music is “unusually complex and alive.” Zhuang Minghao says the US may be more interesting by comparison.
- Figma is the “not-so-good” representative, in quotation marks. Adobe offered roughly $20B to acquire it in 2022; the deal fell through, Adobe paid a large breakup fee and Figma went public independently. Its market cap rose from $10-20B after listing to a peak that may have approached $70B, then fell back to roughly $15B—down more than 88% from the peak and probably still below the IPO offer price—even though revenue growth should have topped 40% for 3 consecutive quarters and may still be accelerating, from roughly 42% and 43% to 45% and 46%.
- Canva has still not gone public. Its secondary-market valuation fell from a peak of “if I remember correctly” about $50B to roughly $32B. The more important signal is that its 2026 revenue-growth expectation was cut from 30% to 20%. It had originally been expected to list in 2026, but rumors now suggest a delay to 2027 or a potential acquisition.
- On the positive-buff side, even as Gen-4.2 and Veo 3 appear to have “obliterated everything,” the top 10 video-model companies have raised money at a furious pace for more than 6 months. Aishi Technology has begun applying for a Hong Kong listing, and a Hong Kong-listed video-model company may soon appear. PixVerse, Hailuo, Sora and Vidu are all worth discussing. The video-Agent field also includes HeyGen, Lifelike, OiiOii and Higgsfield, the US company that recently raised financing. Over the past 1-2 years, these companies have built substantial ARR, user bases and brand influence, while fundraising has moved quickly. In 3D, Meshy and Tripo raised large rounds in the first half, and Zhuang Minghao expects them to consider IPOs soon.
3. What Is the Market Actually Afraid Of? Four Structural Concerns
- The first fear is that models absorb single-point features: background removal, outpainting, background replacement, retouching and templates all “seem to have become default capabilities” for today’s multimodal models. The second is that costs come before revenue. One of the biggest differences between AI and mobile internet is the absence of diminishing marginal costs and network effects. Applications have the lowest gross margins; token, compute, operating, brand, model-call and product-rebuild costs all arrive first, while monetization comes later—and competition may push it even further out.
- The third fear is that experimentation is not retention. “A one-off viral generation may drive sharing…but it is very difficult to get users to use it repeatedly, in batches or as a long-term subscription.” The same applies across multimodality, social companionship and various Agent categories. Taking a somewhat more cynical view, he guesses this may be why everyone is increasingly telling a “bigger story”: traditional activity and retention metrics “really are not that suitable for the state of AI applications at this point.”
- The fourth fear is that growth is not realization. Mature applications must prove that spending on tokens, compute, operations and growth produces not just revenue but profit. In the mobile-internet era, network effects, brands and related factors could generate profit, cash flow and even stronger pricing power. For application companies in the AI era, “at least at this point, that is somewhat difficult.” When the dynamic will reverse—or in which pockets it will reverse—“we do not know either.”
4. What to Watch: From Features to Delivery
- He runs through the industry’s current buzzwords: accumulated user data; memory and context in the broad sense; use-case deployment; end-to-end workflow control, which became more important after Cursor’s rise; controllability, usability and deliverability; collaboration from individuals to teams; and an increasingly important AI brand voice. They all sound like “correct nonsense,” but they can point companies toward their next strategic choices.
- Meitu is a live example of this shift. Its original photo app was “one edit, one effect, one button and done”—a single-point feature. Today it turns a single product image for an ecommerce merchant into a complete asset package, covering store design, paid-acquisition materials and video-production assets. For professional content creators, it links scripts, characters, scenes and the final cut into an Agent-style workflow, allowing one piece of software to handle everything from the initial source to final delivery. Moving from features to delivery has become the product paradigm for AI applications.
5. The Meitu Case: a Legacy Base for Cold Starts, Peach Hides Skill Behind “Learn to Edit Like Him”
- Meitu’s dual identity is the heart of the discussion. On one side is the mature mobile-internet base of Meitu Xiuxiu, BeautyCam and other established apps: a huge user base, strong brand awareness and global distribution that can give new projects a cold start. Zhuang Minghao said his own company has taken a similar approach in its new AI experiments. On the other side is an emerging AI-native matrix of productivity, Agent, subscription, To B and vertical-workflow products covering use cases such as short dramas and MVs. The front end spans image, video and design; the middle layer is Meitu Hub alongside the ZCOOL design community; and the foundation layer is Meitu’s own MiracleVision large model.
- The product he finds most interesting is Peach, whose core feature is “Learn to Edit Like Him.” Photographers and cosers in the cosplay community have their own retouching Skills, but the product does not expose a Skill marketplace or MCP concepts to ordinary users. “For an ordinary user—for a young woman who simply wants to make her photo look better—does she need to know what a Skill is?” The result is a minimal, easy-to-understand path from feature to delivery: users know exactly what they will receive end to end and what it will cost them.
- One counterintuitive rule is worth remembering. The opening line of 郭星’s recent essay “Meitu Survives Like a Startup” is: “Mature legacy products cannot funnel traffic to new products.” This is the decision or rule Meitu ultimately adopted. Zhuang Minghao believes that, on reflection, it becomes increasingly important to everything discussed today.
6. The Real Value of Applications: Internalize Complexity Instead of Dumping It on Users
- His direct criticism of the current popular approach is that asking users to choose a model is “a detour—or, put differently, an expedient compromise accepted at this point in time.” “Users do not need to care which model is being called. The only thing they need to care about is the result: whether the application can deliver a usable piece of work on time.” Looking ahead, that choice should not be the user’s burden.
- Applications need to handle 4 steps. First, choose the model by routing different vision models and tools according to the task and use case. Second, decompose the task by breaking a vague request into elements such as the script, shots, characters, dimensions and file format. Third, control the result by keeping people, brands, templates and styles consistent while allowing users to adjust the creative process and the underlying structure. Finally, complete delivery by producing a visual asset ready to publish, advertise and reuse.
- This also explains his reference to OpenRouter being taken out, as well as the recent emphasis on routing. Both the model layer and the application layer need to select and organize tools based on the task, rather than handing the complexity directly to the user.
7. Two Earnings Numbers, Four Open Questions and the Bottom Line
- In the first half of 2025, Meitu generated revenue of RMB2.21B (+22.1% YoY) and net profit of RMB635M, up nearly 40% YoY. Productivity apps reached 33M MAUs (+43.5%), 2.35M paid subscribers (+29.8%), a paid penetration rate of about 7.1% and related ARR of RMB620M—“on the verge of joining the $100M ARR club.” Traditional imaging products generated RMB1.77B in revenue, with 18.44M paid users and 6.5% paid penetration; the legacy business also continued to grow. Related ARR for recently launched products including Peach and MV Land had already exceeded $500K.
- Two key metrics from the analyst call corroborate each other: third-party API and token-call costs account for only a single-digit percentage of revenue, while the in-house model generates 96% of final delivered content. “This number surprised me so much—it is so large that it feels almost too large.” That is consistent with the New Lab logic: application companies today also need their own models, benchmarks, evaluation systems and post-training reinforcement-learning environments.
- Four questions remain open. Will productivity, the second curve that accounts for roughly 18% of revenue, become the main curve, or “stop growing after reaching a certain scale”? On unit economics, “everyone knows that growth in token consumption does not equal growth in profit,” so the questions are revenue and cost per unit of compute, token pricing, channel revenue shares and gross margin. On product and incubation structure, more products create more shots on goal, but people, money, tokens, compute, brand and R&D resources are all dispersed. Finally, there is the globalization question shared by all Chinese AI companies: long-term retention, renewals and continued payment from users in high-ARPU markets, because MAU growth does not equal payment or subscription growth.
- The final judgment rejects a binary conclusion: “Models mostly eat features…but they cannot eat result delivery in complex scenarios. I think that is something models cannot eat—far from it.” The companies that survive will organize models, use cases, workflows and user payments into outcomes users are willing to buy repeatedly. For companies with established businesses, the answer is “both/and”: stabilize the mature business, turn the existing base into an advantage for incubating new AI applications and workflows, create a positive feedback loop and keep accumulating scale. “Keep moving forward this way without hesitation—that is my judgment.”