Vol.95 The Video-Model Arms Race, the Agent Breakout, and Big Tech’s Anxiety — Crossover with Jinjibo Finance
Summary
The core competition in video models in 2026 will take place mainly among Chinese players, with Google potentially the only US contender still able to compete head-on. 庄明浩 sees Seedance 2.0, Kling 3.0 (adding “if I remember correctly”) and MiniMax Hailuo as evidence that China is building a cluster of leaders; ByteDance and Kuaishou have no strategic reason to exit, while OpenAI has shut down Sora, Anthropic never entered, and Runway has retreated into a multi-model aggregation platform that “sells shovels.” Training, inference, deployment and usage all burn more cash, making near-term profitability unlikely for the leaders; the real barrier is investment scale and strategic commitment without hesitation.
OpenClaw has ignited not a short-lived product cycle but a systemic shift in AI from L1 chat and L2 reasoning to L3 Agents. An Agent must read context, understand delivery standards, absorb feedback and call databases, permissions, files and even the computer itself; the base model is the engine, the reasoning model is the F1 engine, skills, harnesses and the safety architecture make up the car, and tokens are the fuel. 庄明浩 believes the transition “from being very good at chatting to being very good at getting work done” could run through the end of 2026, while model iteration cycles may compress from 6 months to 3 months, 2 months or even 1 month.
The trade in China’s pure-play model companies is initially a scarcity trade, not a financial one. 庄明浩’s “1%-2% comp” simply applies a percentage of US leaders’ valuations: NVIDIA’s $4.5T implies roughly RMB300B-RMB600B for Chinese GPU companies, while OpenAI’s $840B implies roughly HK$60B to more than HK$100B for model companies; Zhipu and MiniMax have already been pushed to 5%-6% of US leaders, or roughly HK$300B-HK$400B. Kimi’s reported valuation jumped from $6B to $10B on the night Kimi 2.5 launched, and Kimi’s PR materials said revenue in the first 20 days exceeded the full-year figure for the prior year; but once the “left foot stepping on the right foot” mutual bid-up stops working, the 6-month unlock around midyear and the 12-month unlock at year-end or early next year become hard risk points.
The market is indulging “new AI” names such as MiniMax and Zhipu while punishing old money such as Tencent, Alibaba and Microsoft, because AI’s direct revenue contribution at the latter may not even reach 1% of total revenue. The S&P information-technology sector’s PE has fallen from roughly 40x back to 20x, reconverging with non-tech companies; NVIDIA trades at about 17x, while Walmart and Costco trade at about 40x, and Microsoft fell roughly one-third in a quarter even though it is difficult to identify what it did wrong. Alibaba is still burning cash in e-commerce and food delivery, while Tencent can use DeepSeek and OpenClaw to showcase its product and ecosystem capabilities—but that is only a tactical positive, whereas the market wants a more aggressive strategic posture.
Anthropic’s most disruptive advantage is not any single benchmark but “coding plus everything,” which connects general-purpose Agents directly to enterprise budgets. 沈帅波 says Claude-related products may have shipped more than 70 features in 3 months; the boundary between coding Agents such as Claude Code, Codex and Trae and general-purpose Agents has already blurred, while a CIO could cut an existing SaaS budget from 100 to 40 and then keep pushing it toward zero. After a model uncovered security vulnerabilities that had persisted for decades, the cybersecurity ETF fell roughly 7% in a day; the reaction was excessive, but Anthropic’s To B revenue curve is forcing OpenAI—more To C-oriented and dependent on memory and personalization for retention—to rethink its strategy.
China’s model cost curve is shifting from “cheap but nobody pays” to “80-85 points and beginning to monetize for real.” 沈帅波 estimates that Chinese models or services may cost only one-seventh as much as US equivalents, and one-tenth at scale; 庄明浩 says the spending required for Zhipu and MiniMax to reach their current level was also roughly 1% of what leading US model companies spent. Users are starting to buy coding plans, tokens and other compute services, while Seedance 2.0 raised prices 3 times within 1-2 months of launch and still could not meet demand; ecosystem companies such as Tencent and Google can also bundle AI with documents, cloud storage, email, sharing and enterprise accounts, making the payment case broader than a single model feature.
AI hardware is more likely to enhance phones than replace them in the near term, while industry diffusion is running into a two-sided mismatch between physical supply and real demand. Hardware reaching consumers today was often completed in the second half of 2025 and initiated in 2024, leaving the underlying intelligence structurally 1 generation behind; return rates for some AI toys may exceed 50%, Chinese glasses makers may still be shipping at the tens-of-thousands level versus Meta’s million-unit scale, and the category remains far from the “early majority.” On the other side, the show says storage capacity is sold through 2027-2028, more than half of US data-center projects in 2026 have been delayed, and copper transmission is beginning to require a shift to fiber and optical communications—while more than 80% of the world has never spoken a single sentence to AI. One side says the story has just begun; the other has already hit hard limits. That gap is the stress test for the current narrative.
Deep dive
1. The main theater for video models is shifting to competition among Chinese players
庄明浩 sees Seedance 2.0 after the Lunar New Year as a watershed: Chinese companies are now widely recognized as being at the frontier of video models, and this is not an isolated lead but “a whole group of companies standing very close to the front.”
Kling 3 was already available before Seedance 2.0, with MiniMax Hailuo alongside it. Even in such a crowded field, startups such as PixVerse can still raise RMB1B-RMB2B in a single round, while some teams that have yet to show much product already carry $1B valuations.
Of the US “Big Three,” Anthropic has never entered video, while OpenAI has shut down Sora and is viewed by 庄明浩 as having exited. In the near term, Google may be the only player still willing to stay in the game. Runway has gone from a pioneer developing its own models to a platform that connects to everyone else’s models—effectively selling shovels.
Video is more expensive than text across training, inference, deployment and usage, and Google, ByteDance and Kuaishou are unlikely to think about profitability in the short term. Kuaishou will not abandon its home turf, while ByteDance already has a lead; the stability of that investment will continue to raise the bar for independents such as Runway and Pika.
2. OpenClaw turns AI from a chatbot into a car
沈帅波 observes that the previous year’s mainstream vocabulary was basically just DeepSeek, while after the 2026 Lunar New Year the discussion has fragmented into “lobsters,” skills, tokens and harnesses. It is as if the menu has gone one level deeper, with the public beginning to understand that AI has internal layers.
庄明浩 uses the L1-L5 framework proposed by OpenAI: L1 is chat, L2 is the reasoning layer represented by DeepSeek and o1, and L3 is where Agents begin. OpenClaw’s significance is that it gives Chinese users a more intuitive view of AI as something that can execute tasks rather than merely answer questions.
His central metaphor is that the base model is an ordinary engine and the reasoning model an F1 engine, but “an engine alone is not enough—you need a car.” The cockpit, steering wheel, brakes, wheels and drivetrain correspond to context, permissions, safety, tools and workflows, while tokens are the fuel.
A real Agent must know the desired result, acceptance criteria and feedback loop, while also calling databases, files, permissions and even the entire computer. Compute consumption therefore rises geometrically. OpenClaw’s heat may fade, but the line from “being very good at chatting to being very good at getting work done” may continue through the end of 2026.
3. The “1%-2% comp” sets a valuation starting point, not a fundamental valuation
At the end of last year, 庄明浩 proposed a “1%-2% rule” for shared China-US themes such as large models, GPUs, commercial space and embodied intelligence: Chinese leaders are initially priced at 1%-2% of US leaders’ valuations, without regard to revenue, losses or business model.
At NVIDIA’s $4.5T valuation, 1% equals $45B, or roughly RMB300B, while 2% is roughly RMB600B. Domestic GPU leaders such as Cambricon were trading around that range at the time.
At OpenAI’s $840B valuation, 1% equals $8.4B, or roughly HK$60B, while 2% is more than HK$100B—close to the valuation range at which Zhipu and MiniMax submitted their listing applications.
The problem is that the two companies have since risen to 5%-6% of US leaders, or roughly HK$300B-HK$400B. 庄明浩 argues that the gap between Chinese and US internet leaders is roughly 10x, so anything below 10% still “sounds reasonable.” But this remains a blunt comp, with none of the classic financial metrics built in.
4. Revenue surges let the primary market realize sentiment faster than the secondary market
庄明浩 notes that the publicly available financial data for Kimi and MiniMax still reflects 2025 revenue. He expects demand from OpenClaw in Q1 2026 to drive a full-year revenue jump that is not merely several-fold but could reach a double-digit multiple.
Kimi 2.5 moved rapidly up various rankings on the night of its release. 沈帅波 says Kimi was raising money at the time, with its reported valuation jumping from $6B to $10B: the original valuation quota was cut in half, and investors’ remaining commitments were priced at the new valuation. Kimi’s PR materials said, “Revenue in the first 20 days exceeded the full-year figure for last year.” Even with zero subsequent growth, that would imply full-year revenue of roughly 18x the prior year.
The awkward position for listed companies was that they still could not issue new shares at the time, so a rising stock price did not immediately replenish cash. Unlisted companies such as StepFun, by contrast, could instantly raise money at a higher valuation. 庄明浩 relays a friend’s joke: “Over there, they’re enjoying a fake stock-price rally; over here, they’re raising real cash.”
5. The self-reinforcing model-stock trade will face its first hard test at lockup expiry
After GLM-5.1 launched, Zhipu could rise roughly 15%; once Zhipu rose 15%, MiniMax seemed to deserve an 8%-9% move, and when MiniMax released another version, it pushed Zhipu higher in turn. “It went up by stepping on its left foot with its right”—every step could be supported by a plausible narrative.
The two companies listed around New Year’s Day. The first cornerstone investors unlock after 6 months, and everyone unlocks after 12 months. 庄明浩 recalls that SenseTime fell more than 50% on its unlock date, “if I remember correctly.” He stresses that this is not a prediction: new models, revenue or another OpenClaw-style event could support the stocks at the time—or fail to.
6. The market indulges new AI while demanding immediate revenue proof from old money
沈帅波 believes the market is highly tolerant of Zhipu and MiniMax but excessively harsh on Alibaba and Tencent. 庄明浩’s response is that AI’s direct revenue contribution at large platforms may “not even reach 1%,” making it impossible to define the platforms as New AI companies in their entirety.
His S&P 500 sector split showed information-technology PE expanding to roughly 40x with the AI boom while traditional companies remained around 20x. Tech PE has now fallen back to 20x, with the two lines reconverging as if “nothing had happened.”
The valuation mismatch is visible in NVIDIA at roughly 17x versus Walmart and Costco at roughly 40x. Microsoft fell about one-third in a single quarter, potentially its worst quarterly performance since the 1990s, yet it still has cloud, Office and gaming, making it difficult to say that it “did something wrong.”
Software and SaaS companies can demonstrate pockets of productivity gains from AI, but when the dominant sentiment is negative, tactical improvements are not enough to reverse valuations. Geopolitics and index weights further amplify volatility; stock prices are not answering only whether the AI product works well.
7. ByteDance, Alibaba and Tencent face 3 different constraints
ByteDance is not listed and therefore does not have to face a quarterly earnings trial in public. Video is also its home battlefield, allowing it to keep investing heavily. 庄明浩 says the benefit of being private is that it does not have to operate under the constraints of every quarterly disclosure.
In the US, AI investment will need to continue at least through 2030 before anyone can see an interim result. Alibaba’s balance sheet is not as deep as assumed: it is still competing and burning cash in e-commerce and food delivery. Whether the previous food-delivery war was worth fighting remains an open question.
Tencent’s problem is not a lack of capability; Hunyuan, organizational changes and 姚舜禹’s arrival all need time. DeepSeek and OpenClaw each gave it a tactical window, and business groups including Cloud, CSIG and PCG quickly began racing one another, producing “personal shrimp, enterprise shrimp and cloud shrimp.” But this has not eliminated strategic concerns about the base model.
Tencent’s stock fell below HK$500 before recovering. 庄明浩 compares it with Apple: the balance sheet is so strong that investors can afford to watch and wait with some latitude, but the market wants a more aggressive posture. Apple, too, set sales records while facing constant criticism; when other phones raised prices because of memory costs, that actually improved the relative value of Apple’s hardware.
8. Q2 2026 remains a joint acceleration of models and harnesses
庄明浩 says new models from 2 leading US companies may already be complete but have not launched because of capability overhang, safety and controllability concerns. He also retains the view that “there is some marketing involved.” The stronger the engine, the more stable, safe and controllable the external chassis needs to be.
There is no standard answer for a harness. The goal is always to let the Agent run autonomously, correct itself and connect to an existing ecosystem or create new tools, but the priority and acceptance criteria for each step are still being tested. Model companies also have to build their own cars because these capabilities will feed back into training the next generation of models.
Chinese companies will expand outward along a similar pattern, but with different emphases: Kimi and Zhipu lean toward Agents and reasoning, MiniMax emphasizes multimodal integration, and StepFun continues to connect hardware and embodied intelligence to the physical world. What used to be a 6-month cycle may now produce a new wave every 3 months, then compress further to 2 months or even 1 month.
9. Anthropic has crossed the gap from skepticism to a scramble for access
沈帅波 summarizes Anthropic after the Lunar New Year as “unstoppable”: he says Claude-related products may have added more than 70 features in the past 3 months, almost one new release every day, without slowing the pace of model evolution.
Anthropic’s strategy can be reduced to “coding plus everything.” When people who cannot code use Claude Code, OpenAI Codex or ByteDance’s Trae, they are still simply describing a task and receiving the result; they do not need to understand the underlying code. The boundary between coding Agents and general-purpose Agents has therefore nearly disappeared.
The adoption curve is not arrogance, skepticism and gradual acceptance, but a rapid move “from skepticism to a scramble.” Once enterprises see that others are already using the tools, compute scarcity can make even the most expensive Claude a sought-after product, with demand released nonlinearly after model capability crosses the threshold.
10. “Coding plus everything” is rewriting SaaS budgets directly
Enterprises historically handed budgets to software and SaaS companies through subscriptions, corporate cards and CIO procurement. Once a CIO realizes that internal Agents can rebuild the workflow, the budget may not fall from 100 to 90 or 80; it may be “cut instantly from 100% to 40%, and then cut to zero.”
沈帅波 says the revenue figures Anthropic has disclosed in recent months have nearly doubled month after month. The market keeps mapping its new capabilities onto ServiceNow, cybersecurity and other enterprise-software sectors, creating a single causal chain: Anthropic revenue rises while SaaS valuations fall.
One model demonstration uncovered vulnerabilities that had been hidden for decades in software known for its security. The cybersecurity ETF fell roughly 7% across the board that evening. 沈帅波 acknowledges that selling without distinguishing between business lines was “certainly excessively pessimistic,” but the market is repricing old budgets one cut at a time.
OpenAI was originally more To C-oriented, building revenue around users, retention, time spent, conversion, memory and personalization, while individual willingness to pay naturally has a ceiling. Anthropic’s To B curve is steeper, forcing OpenAI to narrow its scope and reset priorities; the two are engaged in a rare “knife fight” inside a mature, major market.
11. China’s cost advantage is beginning to translate into real payments
沈帅波 estimates that Chinese models or services may cost one-seventh as much as US equivalents, and one-tenth at scale. 庄明浩 adds that the money Zhipu and MiniMax spent to reach their current level was also roughly 1% of the spending by leading US companies, yet they can deliver an “80-to-85-point” product.
His reservation is that model capability is still rising rapidly; the market has not reached a stable phase in which the US pauses and China catches up through refinement. The next few years will still be a “frantic chase.” Video may be the first area where quality and price converge with or even surpass the US, because China has already developed several strong competitors.
More important, Chinese users are moving from “zero or negative” willingness to pay into positive territory. Some are buying coding plans, tokens and other compute services, while government subsidies also correspond to real spending. Seedance 2.0 raised prices 3 times within 1-2 months of its official launch and still could not meet demand, overturning the negative assumption that domestic AI can only be free.
12. Entrances and ecosystems lock in payment better than standalone model capability
沈帅波 uses his own payment behavior to illustrate the path: he used to dislike WPS and would not pay for it, but was willing to pay more than RMB100 after AI was added. He previously paid for Youdao Cloud storage and now buys Tencent Docs because the same payment covers storage, AI and connections across the Tencent ecosystem.
Tencent has a broad ecosystem in China, while Google has Gmail, Drive, Docs and To B products in the US. 庄明浩 argues that AI is not an isolated feature but a connective layer running through the stack, so the relevant question is how much it adds to the ecosystem as a whole.
After the trial period, individuals will keep only 1 or 2 tools. Once an enterprise standardizes on Feishu or DingTalk, its accounts, memories, operating habits and sharing practices become embedded as well. One-click prompt and data migration show that vendors are not fighting for a single call; they are fighting over which ecosystem users ultimately converge on.
ChatGPT’s strategic adjustment is to bring chat, Codex, a browser and other capabilities into an all-in-one product. Anthropic is extending from the model to its own “car,” launching an Agent called Manus that charges by runtime. Model fees, token fees and Agent execution fees are beginning to stack.
13. Hardware’s annual cycle cannot keep pace with software’s monthly iteration
沈帅波’s experience is that an AI toy appears advanced when it first arrives, then exposes its weak understanding and memory of the physical world 2 days later, still relying mainly on fixed actions that trigger preset responses. He believes the root problem is the mismatch between software updated monthly and hardware developed annually.
Hardware reaching consumers today was probably completed in the second half of 2025 and initiated in 2024, while DeepSeek R1 was not released until the end of January 2025. Even with OTA updates, the underlying software, hardware and product form fixed at the project’s inception cannot be completely rebuilt.
Toy companies can only start by narrowing the feasible use cases: plush or hard-shell, with or without a camera, mobile or stationary, standalone or phone-connected, focused on companionship or learning. Each generation may add only a little intelligence, while the news cycle is already discussing the next generation of Agents by the time the product launches.
As a result, return rates for some AI hardware are “about the same as women’s clothing,” potentially exceeding 50%. Add live-streaming channel fees, storage and chip costs, and next-generation R&D, and the commercial barrier is far higher than a demo suggests. Hardware founders must also carry the pressure of supply chains, employees and historical investors.
14. AI glasses have not crossed the chasm; the phone remains the strongest hub
沈帅波’s view is that every new hardware category claims it will replace the phone but in practice enhances it. A recording badge improves collection efficiency, but the first step is still to send the content back to the phone; even optimized glasses typically need to connect to one. The phone increasingly resembles a more powerful router and hub.
A carmaker’s glasses get the weight, camera, voice and large-model connectivity “right” and can link with the vehicle, making them attractive enough to owners. But outside that brand ecosystem, ordinary users still lack a compelling reason to buy them, and this is not a problem unique to one manufacturer.
The show compares Meta’s glasses with Chinese manufacturers by order of magnitude: Meta is at the million-unit level, while Chinese companies may still be at the tens-of-thousands level, a gap of roughly 2 zeros. Domestic products may serve a niche, but they have not reached the “early majority,” much less reproduced Google Glass’s old dream of becoming a universal entry point.
The show mentions a highly ranked glasses manufacturer with 9-10 years of operating history whose cumulative sales are only several hundred thousand units. Its cash flow has been under pressure since founding, and it could even break without an IPO and new financing. 沈帅波 still leaves a sliver of hope: “Maybe we can get excited one more time—who knows?”
15. AI is both reinforcing old industries and hitting narrative and physical limits
Lenovo’s record revenue led 沈帅波 to reassess “old businesses”: PC demand, replacement cycles and AI enhancement mean that even a fraction of the company’s scale can exceed the entire size of a hot startup. Capital prefers sexier new stories, he says, but the business performance of old ecosystems after absorbing AI is equally real.
NVIDIA traded sideways around $170-$190 for 8-9 months, despite 2 strong earnings reports during that period. 沈帅波 does not think it can remain range-bound forever: the outcome is either that it “stores up another big move” or heads lower, but he explicitly says he cannot determine the direction.
Hard supply constraints are already appearing: TSMC’s leading-edge process has reached 2nm, reported 2027-2028 memory and hard-drive capacity has already been sold, more than half of US data-center projects in 2026 have been delayed by issues including land approvals, and copper’s transmission rate is insufficient, making a shift to fiber and optical communications necessary.
沈帅波 looks back at people still invested in the Web3 narrative and worries that AI could also “have value and exist, but be unable to carry such a title.” He still believes AI is far bigger and more solid than Web3, but white-collar anxiety, anti-AI protests and rising unemployment are reminders not to mistake emotional amplification for certainty.
16. AI diffusion is highly uneven; 10 years after AlphaGo, everyone is now in the game
What 沈帅波 saw in Europe was a different technology path: credit cards and contactless payments are already mature infrastructure, so locals have little incentive to detour into China-style mobile payments. Americans say they will change the world, Chinese users quickly follow, while Europeans first ask about environmental impact, rules and law. Technology never diffuses in lockstep.
For most ordinary Chinese users, AI is basically Doubao. 庄明浩 says Doubao combines vision and LBS to identify museum buildings, artworks and attractions, effectively “killing” the tour-guide-device business. The parents’ generation uses AI to generate short videos on Douyin; China is particularly good at turning capability into To C entertainment and everyday experiences.
At the other end of the spectrum, a business-group president at a Fortune 500 company—roughly in the top 100—did not even know Anthropic, Claude or Gemini. 沈帅波 later realized that the executive already had multiple assistants handling retrieval, organization and reporting. Those assistants were effectively his Agents, so he had little reason to learn AI personally.
On the 10th anniversary of AlphaGo’s victory over 李世石, the show views Grok playing T1 and Faker as another historical replay. This time, AI must also see the screen, use a mouse and keyboard, and handle complex coordination and split-second decisions. “At first you hear the tune without understanding it; hear it again and you are already inside it.” Ten years ago, 黄仁勋 repeatedly said GPUs served more than gaming, but the idea was ignored for years because “people can believe only in the world they can see.”
Verification Notes
- In the valuation section, A says that Kimi and MiniMax had financial statements and were listed early in the year; later passages describe Kimi and StepFun among the unlisted companies. The source pipeline does not harmonize Kimi’s listing status, and the claims are retained separately.