Primary-Market View of 2026 AI's First Half: FOMO, Data Loops, IPOs
Primary-Market View of 2026 AI's First Half: FOMO, Data Loops, IPOs
Summary
- Two Fengrui technology investors’ review of the 2026 first-half primary market comes down to 3 words: FOMO, divergence, and front-loaded valuations. In hot sectors, “a $1B valuation can be born rocket-like within just a few months”; one deal may be closing, another in term-sheet talks, and a third already finalizing its structure. But the heat is highly concentrated, with most sectors moving sideways—an almost perfect mirror image of the extreme divergence in technology stocks on the secondary market.
- OpenClaw (heard on the audio as “Open Cloud,” inferred from the “raising lobsters” context) was the first half’s most globally influential event. It turned agent from investor jargon into a mass-market consensus around “hiring digital employees,” much as DeepSeek broke into the mainstream. Once the hype receded, what remained was real usage and a new paradigm: harness engineering, task-trace data, and agent RL. The core investment question is whether vertical agents’ trajectory data might not flow back to the foundation models and whether “a mechanism that can be accumulated may form at the agent layer”—the reason agent startups may avoid being swallowed by model companies. Cursor is the counterexample: too close to the model, “like a star expanding at the end of its life—you’re a planet too close to it and get swallowed instantly.”
- Anthropic is the standout company on the model side. Going all-in on coding was an “excellent strategic decision,” taking the company from follower to leader; reported ARR has reached $16B, an IPO this year could imply a $1T valuation, and the company may break even by year-end. The evaluation regime has shifted from exam leaderboards to SWE-Bench-style task-delivery rankings: “The model narrative has moved from ‘I can sit an exam’ to ‘I need to deliver real agent value.’” Fable 5/Mithos is reportedly not being offered externally and may be subject to a ban; 严千行 believes the story contains “an element of smoke and mirrors, and an element of hunger marketing.”
- The US-China gap may be narrowing to “a few months.” Zhipu’s GMM 5.2 is viewed as approaching GPT-5.5 and the latest Claude Opus on coding, while its market cap has crossed HK$1T; but the infrastructure gap represented by clusters with hundreds of thousands of cards cannot be solved in the short term, leaving China with a “do more with less” playbook. Two clear positions are lightweight models—MiniMax 2.5 became OpenClaw’s official recommendation, while Kimi and MiniMax appeared in Jensen Huang’s GTC deck—and domestic strength in video generation: Seedance’s revenue run rate has reached $2B as Sora shuts down, with the decisive factors being video-platform data and inference costs.
- The infrastructure narrative is shifting from “do we have enough compute?” to “is the system operating efficiently?” Nvidia has traded sideways between $180 and $220, while optical interconnect and memory names have delivered tenfold gains—“it’s starting to look like crypto.” The transmission wall and memory wall are forcing optics to displace copper and HBM to become ubiquitous; inference will account for more than 90% of demand, and memory inflation has already reached MacBooks. The entire system can be reduced to one equation: “power in, tokens out.”
- World models were the first half’s clearest FOMO trade. Dozens of companies and 4 competing routes—video generation, 李飞飞’s 3D approach, LeCun’s JPA, and WAM—have yet to converge, while financing “went bang-bang-bang straight up.” The embodied-AI reset is that VLA and world models are complementary rather than substitutes, like tennis: instinctive reaction versus thinking before moving. North America’s latest narrative is a goal-driven systems approach, “Beyond VLA and World Models.”
- The IPO wave is changing how value transmits between the primary and secondary markets. Unitree is preparing to list on the STAR Market, Zhipu and MiniMax are already public, and OpenAI and Anthropic may list this year—the essence is “capitalizing and realizing ecosystem-position value,” while replenishing ammunition. The risk is that “if the leading names cannot support market expectations, they will in turn suppress valuation expectations across our primary market,” with a primary-secondary inversion triggering a “major cooling-off period.”
- The bubble will eventually break—calling that “a correct but useless statement”—but capital will consolidate after the break. Research companies that can deliver milestones consistently will continue to raise money, while “companies with no results beyond storytelling will not”; deployment-focused capital will move toward businesses with real orders and robust commercial models. 严千行’s key second-half variable is a genuine data loop: embodied AI rolling continuously from data to deployment, agents making decisions and learning on their own, because “skills do not constitute a business model.”
Deep dive
1. 李丰 Opens: Put the Bubble Question on the Table Amid the Shock
- Fengrui partner 李丰 set the tone: over the past week, AI equities—chips and the memory supply chain—“have seen enormous volatility.” Combined with Meta’s plans for its own compute infrastructure, the market is asking whether compute infrastructure has reached a cyclical high or the point where the bubble starts to break.
- He handed the microphone to 2 technology-investing colleagues while stressing that the fund favors “open and transparent internal discussion” and that each view represents the speaker alone. His own view on the stage of the AI and token cycle was laid out clearly in the macro discussion 2 episodes earlier; “the discussion below also contains some differing opinions.”
2. 3 Keywords: FOMO, Divergence, Front-Loaded Valuations
- 严千行 spoke with nearly 150–200 AI projects during the first half, and 刘鹏奇 had a similar experience. Their 3-word summary starts with FOMO: hot areas are producing a flood of unfamiliar projects, and rapidly rising valuations are forcing investors to “look at every single project.” The internal pressure becomes: “Did we invest today? Do we need to invest more?”
- The second word is divergence: “Apart from a few exceptionally hot sectors where financing keeps accelerating, most sectors remain relatively calm”—exactly the same structure as the secondary market, where technology stocks surge while non-hot names grind lower.
- The third is front-loaded valuations. The old early-stage rule of assigning a reasonable valuation before milestones has been “completely broken”; “a $1B valuation can be born rocket-like within just a few months.” 刘鹏奇 supplied the live footnote: one project’s valuation changed within a week, with “one round closing, another negotiating terms, and a third already finalizing its structure.”
- The underlying anxiety is not simply fear of making a bad investment. 严千行 admits that “saying you’re not anxious would be unrealistic,” but the bigger fear is being left behind by the AI era: whenever a new model or buzzword appears, the worry is, “How did I only see this news several days later?” 刘鹏奇’s answer is to separate projects into value and opportunity buckets, so price-value mismatches do not distort decisions.
3. A Double Return: New-Lab Research Startups and “FOMO Is Useless”
- The first shift is toward research-driven sectors: the US-style new lab uses the organizational form of a commercial company to accelerate long-term research. Investors are less concerned with whether the founder is a professor; the question is now, “Can you be the best research producer in this industry?”
- The second is an awakening at the application layer. Reviewing the FOMO timeline, 严千行 says investors chased AI applications in the first half of last year amid the suspected Manus craze, then chased AI hardware in the second half amid Insta360’s IPO. “But after returning to the first half of this year, everyone realized one thing: FOMO is useless.” The market has settled back on better product definers: agents must deliver task value, close the task loop, and build their own data-flywheel moat; hardware must answer whether consumers are genuinely interested and have a need.
- These apparently opposite directions “are in fact both a return to value in their respective startup formats.” That, in his view, was the first half’s biggest trend.
4. The Nail Is Doing the Hammering: Coding Tools Rewrite the Founder Profile
- Since Codex and the suspected release of Claude Code, founders no longer have to be engineers. Strong product managers and designers can “turn themselves into full-stack product managers.” The classic joke from 10 years ago—“I have an idea and an angel round; all I need is a CTO”—no longer holds.
- 严千行’s observation from offline hackathons is that the best products often come from lawyers, stay-at-home mothers, and non-IT workers. 刘鹏奇 summarizes the old model as “taking a hammer to look for a nail”; now the nail starts the company. 严 responds: “The hammer is just buying a coding plan. I know where the nail is, so I’ll hammer it myself.” Fengrui has encountered several strong products this year from teams without an algorithm background.
5. OpenClaw Breaks Out: From Insider Jargon to “Raising Lobsters”
- 严千行 calls OpenClaw (heard on the audio as “Open Cloud/Open C O L,” inferred from the “lobster/raising lobsters” context) “the most globally influential event” of the first half. It transformed the agent concept discussed by investors and founders into a mass-market consensus: “I want to hire a digital employee.” The conceptual shock resembles DeepSeek’s breakout last year.
- The breakout mechanism was straightforward. Deploying an agent used to be a hassle, and most people had never used Codex or Claude Code. OpenClaw was the first to let ordinary users deploy quickly on the edge and “raise” an agent through natural-language conversation. Open source and low barriers, backed by Chinese open-source models, then created a social-media feedback loop.
- Another signal appeared in the secondary market, where the joke “Are you standing in the light, or standing there naked?” went viral. “People who do not do AI, do not understand AI, and only trade have today become the group that understands AI best and spends the most time studying it.”
6. After the Tide Goes Out: Foam on Top, Rich Beer Below
- The fading heat may also reflect open source’s dual nature: low barriers create rapid adoption, but weak safety mechanisms and an unfinished user experience send most casual users elsewhere. 严千行’s conclusion is that retention matters: “Many people who were attracted in do not actually need agents, but the people who do need agents are genuinely using them steadily right now.”
- His best example is a beauty blogger friend who became a technology blogger. With no technical background, she used OpenClaw to build her own workflow, gather targeted technology information, and automate a content pipeline, then used social platforms to amplify her output. PPT creation is another now-standard use case. 严千行 adds: “A lot of the BPs we’ve received this year are obviously made by AI.”
- His closing metaphor is worth preserving: “After a wave of excitement, there is foam and noise, but once it settles, you find that it has genuinely blown off the foam from the beer, leaving rich beer underneath.”
7. Harness Engineering: Let the Model Explore Instead of Hard-Coding the Path
- The fundamental jump from chatbot to agent is a closed-loop system of environment setup, decision and execution, and feedback; the model is only the brain. This year’s keyword is harness engineering: define the environment and feedback channels, then let the model run the workflow itself. “You do not need a person to hard-code every path. AI can explore better decisions on its own.”
- Why was this not a theme last year? 严千行’s explanation is that the model’s exploration within a given environment “was not necessarily as effective as a workflow you explicitly defined.” Now tool use and long-horizon decision-making are stronger, and “what it figures out will definitely be better than what we prescribe.”
- His evolutionary analogy is one of the episode’s central images: once the environment is built, the agent ecosystem is “like the Cambrian explosion.” Each iteration is not necessarily positive, but with a sufficiently good environment and abundant resources, participants randomly evolve inside it, and the result is an explosion of forms.
8. Task Trajectories: When the Agent Data Flywheel Can Work
- 刘鹏奇 raised the old question: foundation models still lack a genuine user-data flywheel. 严千行’s answer is that the data has changed shape. In the chatbot era, the key asset was contextual text; in the agent era, task trajectories are the core data, analogous to the real-machine task data prized in embodied AI. Combined with a harness environment, these trajectories naturally drive the agent RL that has become popular this year.
- He does not avoid the key uncertainty: can trajectory data flow back to the foundation-model companies? Claude Code’s coding trajectories certainly feed Anthropic, which is why “it can do better than others.” But agents will be widely distributed across verticals, and he believes trajectory data may not be open to foundation models: “I think some mechanism that can be accumulated may form at the agent layer. That is what creates room for agent startups.”
- This directly answers the bearish view that stronger models will swallow every boundary. Task trajectories and vertical data may circulate within their own systems, allowing agents to build their own data flywheels and avoid being swallowed by large models. Whoever controls those assets will hold the strategically contested ground.
9. 3 Camps: Stars Swallow, IM Platforms Absorb Workflows, Lobster Boxes Remain Unproven
- Who wins? 严千行 first warns that “trying to say today who has the higher win rate is pretty close to fortune-telling,” but the 3 camps want different things. Model companies are taking the general-purpose use cases closest to the model and capable of feeding it back—deep search, deep research, and coding. Cursor is the tragic footnote: it was acquired after going from the largest source of token consumption for Claude Code to being killed by Claude’s own agent product—“like a star expanding at the end of its life: a planet too close to it gets swallowed instantly.”
- IM and office-productivity entry points have ready-made distribution ecosystems and accumulated workflows. Standard digital processes such as invoice reimbursement raise the question: can an agent learn them quickly? In other words, SaaS-era workflows can become agents directly.
- For local “lobster machines,” he puts a question mark on the idea: “When the hype first started, the thing people fought hardest to buy was the Mac mini; once it cooled, the thing most commonly sold on Xianyu was also the Mac mini.” Unless usage is clearly frequent and privacy requirements are high, a personal computer or cloud server is enough.
- Derivative forms of edge AI hardware—compute boxes, edge routers, and edge hard drives—are still supplements driven by AI use cases. That leads to the next question: who owns the data?
10. Trading Privacy for Capability: The Real Bottleneck to Releasing Enterprise Tokens
- 严千行 divides users into 2 layers. Historical experience from the consumer internet suggests C-end users “care more about efficiency and experience than privacy, and may even give up some privacy”; much of their data is already held by platforms, and they continue to enjoy the services. Pro C and B-end users are completely different: data assets are the core of a company’s operations, and handing them to a foundation model in exchange for efficiency is a losing trade.
- He cites conversations with friends at Alibaba Cloud as evidence: C-end customers generate a reasonable volume of token calls, but large B-end customers’ token usage has never taken off, wildly out of proportion to their cloud-computing consumption. They cannot release model capability and tokens into business scenarios, so usage remains limited to relatively shallow productivity tools.
- 严千行’s deduction is that lawyers cannot disclose client files and engineers cannot disclose formulas: “Either deploy privately or do not do it.” Once privacy and connectivity are solved, these scenarios become the next entry point for AI. 刘鹏奇 confirms this is exactly what Fengrui is watching: connecting the best model capabilities to local knowledge bases, with permissioning and privacy protection.
11. Enterprise Deployment: FD, “More Volume at the Same Price,” and Assigning Hallucination Liability
- Both investors have experience with SaaS and FinTech and are “both fond of and frustrated by” large B-end customers: they have money, high standards, and complicated problems. The immediate reality of selling into large enterprises is “more volume at the same price”—customers believe models are open source and do not want to pay an AI-service premium, while domestic price wars are particularly aggressive.
- The overseas FD model—AI deployment engineers working as outsourced implementation staff—is hot, but 刘鹏奇 warns that once these engineers reach the customer’s underlying databases, they face the same old B-end problem seen in China. The more fundamental obstacle is hallucination liability: “If I run a risk-control decision and hand it completely to AI, who bears the loss caused by a hallucination—the model or the employee? That allocation is like autonomous driving; there is currently no way to split it.” This is why AI remains in productivity scenarios rather than entering core business processes.
- The bottleneck itself creates opportunity: standardizing and connecting internal enterprise data with foundation models, building business graphs, and controlling hallucinations. “Some of this may still require academic problems to be solved, but that is also where the opportunity lies.”
12. The Agent Paradox and Token Economy: The Pricing Anchor Has Not Been Found
- The paradox is explicit: “If Claude can do everything, why start an agent company?” 严千行’s answer is the classic scale argument: the market always extracts common needs and turns them into products refined by the best operators. He wanted to use Claude Code to build an AI RSS summarizer, then found an existing product that worked well; paying “a dozen or so dollars a month” was better than burning hundreds of dollars in tokens to build it himself. “Once a product’s common needs are clear and it can scale, buying is cheaper than building.”
- 刘鹏奇 adds a more macro caution: much of the DIY activity is driven by the satisfaction of proving, “We are only missing an engineer, but even without one I got this done.” Yet “a lot of the code developed by C-end users is really garbage code sitting there,” making the real economic value per token uncertain. Whether tokens should be priced on cost or value “is still taking shape”; the subscription model cannot be linearly extrapolated to the revenue model of large model companies.
- 严千行 views the recent spike and retreat in token spending as healthy: in the past, “when I was uncertain, I preferred to use the most expensive option.” Now, “do not use a sledgehammer to kill a chicken; a cheaper model works too”—a normal economic fluctuation.
13. Models Are Still Improving: From Exam Rankings to Task Delivery
- 严千行 draws a 2-part distinction. Structural, step-change innovations of the o1 reasoning type—“we really have not seen them for a while”—but individual capabilities continue to improve rapidly. The clearest evidence is the evaluation regime: it has moved from MMLU and math exams to SWE-Bench and Computer Use agent benchmarks. “The model narrative has moved from ‘I can sit an exam’ to ‘I need to deliver real agent value.’”
- 刘鹏奇’s analogy is clean: it is like a student graduating into society. “The knowledge base is basically set; evaluation is no longer an exam score, but how many work tasks you can complete.” AI coding is becoming more useful because the models themselves are getting stronger.
14. Anthropic Aligns Knowing With Doing: The Strategic Dividend of Going All-In on Coding
- The standout company is, without hesitation, Anthropic. Putting all its weight behind coding was an “excellent strategic decision,” taking it “from a follower in the foundation-model race to the industry leader.” It saw earliest that evaluation would shift from exams to task execution; “when Claude Opus 3.5 came out” (his wording), coding tools became useful, creating a positive use-to-improvement feedback loop.
- The numbers are extraordinary: reported ARR has reached $16B, the company could list this year at a projected $1T valuation, and it says inference gross margins may reach break-even by year-end.
- He rejects the idea that coding is a vertical market. Coding contains human logic and decision processes, so investing in coding feeds back into reasoning and chain-of-thought decision-making. “People find it especially rigorous and smart—not naming certain model companies, whose products feel imaginative but unreliable.” His summary line is the episode’s signature: “Learning without thinking leads to confusion; thinking without learning leads to danger”—Anthropic was the earliest and best at aligning knowledge with action, so its “knowing” advanced much faster than everyone else’s.
15. Fable 5/Mithos Ban Rumors: Weapons-Grade Capability or Hunger Marketing?
- Anthropic’s release of Fable 5, also called Mithos, generated a steady stream of headlines this year: claims that it would never be offered externally, a brief opening followed by a takedown, and rumors that the US would ban it. GPT-5.6 was also released but not opened to the public.
- 严千行 is cautious. The capability improvement is real—context length, adaptive iterative reasoning, and integration with agent capabilities produce a noticeably stronger user experience. But the ban narrative “contains some smoke and mirrors, and some hunger marketing”; whether it truly represents a national-competition weapon remains to be seen.
16. China’s Biggest Surprise Is Zhipu: GMM 5.2 and Catching Up on Data
- 严千行’s answer is immediate: “The biggest surprise came from Zhipu”—its stock price, with market cap just crossing HK$1T, and GMM 5.2, which likewise places coding at the center. 刘鹏奇 adds a technical detail: beyond proprietary architectural innovation, 5.2 reportedly introduced process rewards into long-horizon task training.
- The more important structural change is catching up on data. In response to questions about whether Chinese companies are doing extensive distillation—“the ceiling for distillation alone is limited”—domestic giants, after raising capital through IPOs, are increasingly willing to spend serious resources and money on data collection in addition to buying compute and talent. The data-service providers Fengrui has seen this year are making substantial money: any data available is being snapped up.
17. DeepSeek’s Exception: The Competitive Unit Has Changed
- DeepSeek completed a financing round this year after long insisting it would not raise capital. 刘鹏奇 sees both talent and resources behind the move: valuation-based employee incentives matter because “this is the most important asset of this era,” while deeper ties to the domestic compute ecosystem can “not only improve its own capabilities, but bring the entire Chinese industrial ecosystem up with it.”
- 严千行 supplies the harder context. When DeepSeek first emerged, domestic competitors had funding in the tens of billions of RMB, and quantitative funds’ research labs had enough resources to compete. But since last year, “the major platforms started throwing money at it”—ByteDance-linked teams, Alibaba, and Xiaomi are all spending heavily. After MiniMax and Zhipu listed, the ammunition gap widened further. “If DeepSeek continued operating in the old way, it would clearly be making the challenge harder for itself.”
- He remains respectful in his characterization: “Everything is in service of building a better foundation model, not in service of a commercial-deployment objective.”
18. The US-China Gap Is Down to Months; Lightweight Models Have Already Arrived
- 严千行’s layered view is that the US-China gap is narrowing in fast and general-standard models. Many believe GMM 5.2’s coding capability is now very close to GPT-5.5 and the latest Claude Opus. But one clear problem remains unresolved: infrastructure and compute. The US has clusters with hundreds of thousands of cards, making the frontier difficult to catch in the short term; China is better at algorithmic efficiency and “doing more with less.”
- He gives the gap a specific timeline: in 22, “it felt impossible to catch up”; in 23–24, the gap was “1–2 years”; this year, it is “just a few months.” Evidence came from an exchange on X between Zhipu’s CEO and Elon Musk. Musk asked when China would catch up; the answer was: “After 1 quarter, we may catch up next quarter.”
- China has already secured a position in lightweight models. MiniMax 2.5 became the official recommended model when OpenClaw launched. At Nvidia’s GTC in March, when Jensen Huang discussed model tiers, “the names Kimi and MiniMax were both on his PPT.”
- 刘鹏奇 adds a structural point: data’s contribution to model capability will keep rising, and China has become much stronger at filling its data gap. He is watching that angle closely as well.
19. China Gains an Edge in Video Generation: Seedance’s $2B Revenue Run Rate vs. Sora Shutdown
- The competitive picture has unusually reversed: Seedance’s revenue run rate has reached $2B, while overseas offerings have little to show beyond Google Veo 3.0, which still receives recognition. “Sora has already shut down its service.”
- 刘鹏奇 identifies 2 reasons. First is data: ByteDance and Alibaba hold large libraries of existing video data and possess the engineering capability to process it. Second is cost: video models consume enormous numbers of tokens. “Generating the same video with Sora costs many times more than with Seedance; it is hard to sustain a business model at the same price.” Seedance has used architecture and compute more efficiently to bring costs into an acceptable range, and video-related revenue now represents a meaningful share of ByteDance’s overall business.
- 严千行 adds the demand and data loop. China has real consumption of AI short dramas and AI manga dramas, while studios “are always waiting for the latest model improvements to upgrade their production workflows.” The 3 strongest video-model companies—ByteDance, Kuaishou, and Google—each have their own video platforms. They may not directly use platform data, but they know how to process those data assets from an engineering standpoint.
20. Infrastructure Repricing: Nvidia Goes Sideways While the Upstream “Looks Like Crypto”
- The first half packed in enough major events: Nvidia acquired Groq at the end of last year; GTC in March introduced an LPU solution; Cerebras, the integrated-wafer company, went public; and OpenAI taped out its own inference chip just 2 days ago. The sharpest price signal is that Nvidia has traded between $180 and $220 without a major move, while upstream optical-interconnect and memory companies delivered “more than tenfold gains that make people jealous”—“it’s starting to look like crypto.” As the joke goes, “You put your year-end bonus in last year and can earn a full year’s salary this year.”
- 严千行’s core takeaway is that the question has shifted “from whether there is enough compute to whether the entire compute system is operating at maximum efficiency.” The industry is no longer using Nvidia’s best GPGPU for everything; it is building systems that maximize ROI for each use case.
- Inference is the foundation of that logic. If AI truly penetrates every industry, “inference will absolutely account for more than 90% of total demand,” and inference cost will determine the pace and volume of token adoption across society.
21. The Transmission Wall and Memory Wall: Optics Displaces Copper, HBM’s Value Is Unlocked
- The plain-English version of the transmission wall is useful. In a 10,000-node cluster, if transmission cannot keep up, “you have hired 10,000 people, but because the assembly line moves slowly, some of them are slacking off—you have wasted the money.” That is how GPU utilization gets eaten. Optical-module use is therefore surging, board-level connections are moving toward PCB backplanes, and pluggable optical modules are evolving toward NPO/CPO; the long-term direction is “optics in, copper out.”
- The memory wall works the same way: moving data leaves compute units idle, so the industry must raise memory-compute bandwidth. HBM, which was regarded for more than a decade as an expensive, low-value product, has suddenly been deployed across the industry at scale, unlocking its value. HBM’s high cost is also driving tiered-storage demand, pushing SSDs and DRAM higher together.
- The industry’s character is changing. Previously, suppliers developed new technologies and then searched for applications. “Today Nvidia is forcing the upstream supply chain to solve problems and accelerate R&D—AI’s explosive demand is turning a cyclical industry into a growth industry.”
22. Price Increases Reach Everyone: Power In, Tokens Out
- Demand crowding is already visible. An Alibaba Cloud contact says, “In 2026, you’ll have to fight to buy memory.” Apple announced yesterday that MacBooks will become more expensive because of higher memory costs. 严千行 adds a personal footnote: “Why was I so determined to replace the phone I had used for 5 years last year? If I wait until next year, the phone will probably be more expensive because memory prices will have risen too.”
- He reduces the entire compute system to one line: “Power in, tokens out.” The need to convert OPEX into output as efficiently as possible will keep forcing every module to improve—lower power, lower cost, and better maintainability. “This is all about building better AI infrastructure and supplying it to everyone at a lower price. Like foundation models and agents, it is solving the problem underneath: making tokens cheaper, better, and faster to use.”
- One counterintuitive short-term implication, raised by 严千行 and partly accepted by 刘鹏奇, is that cloud infrastructure may absorb resources and push up hardware prices, “suppressing edge demand in the short term.” Consumers may postpone replacing computers as even cheap edge hardware gets more expensive. 刘鹏奇 notes that in 2 years the market will also have to contend with an AI-cycle correction and upstream bubble deflation.
23. Edge, Fog, and Cloud: Tens-of-Billions Models on the Edge and Full-Power Open Source at the Fog Layer
- 刘鹏奇’s forecast is that the C end will have a relatively high edge share. Startups are making it possible to run models with tens of billions of parameters on the edge; simple tasks and some agent tasks could eventually run entirely locally, while coding and long-horizon reasoning are handed to the fog layer or cloud.
- The fog layer is positioned as the carrier of private data: a NAS or all-in-one device that stores personal and enterprise documents, video, and databases. It can “basically run a full-power open-source model locally,” with sensitive tasks sent to the cloud through a de-identification mechanism for collaboration.
- This creates an opening for startups. Since cloud infrastructure will absorb resources in the short term, “edge systems will in the long run favor better functionality with relatively low compute consumption and low hardware cost.” Compressing large-parameter models and enabling them to interact with SSD storage to fit into low-performance hardware is a direction Fengrui is funding, though “how large a market this model can drive still needs to be observed.”
24. World Models: 4 Routes Behind a Super-Hot Buzzword
- There were “probably dozens” of companies claiming a connection to world models in the first half. 刘鹏奇 first strips away the mystique: a world model is not a specific model, but the concept of “intervening in the world externally and predicting how the state will change at the next point in time.” By that definition, “Newton’s laws are actually a world model.” 严千行 adds from his reinforcement-learning background that the middleware in a stimulus-feedback-reward loop is already a world model.
- Route 1 is the video-generation camp. Sora once called itself a world model, “but it was not very accurate—hands went through glass, and people getting into cars passed straight through the window.” Route 2 is 李飞飞’s 3D world model: generating an interactive, operable 3D space rather than frame-by-frame flat images. “Whether it can achieve real physical interaction, we have not seen concrete progress yet.”
- Route 3 is LeCun’s JPA, led by “a highly vocal standard-bearer against LLMs and next-token prediction.” 严千行’s microphone analogy captures the idea: next-token prediction must predict every pixel in the scene, while JPA “only needs to know that the object in front is a microphone.” It abstracts the scene into conceptual representations in latent space. “Whether it can actually be built, I do not know today; it is challenging one of the hardest problems in the world.” Route 4 is WAM, or world action model, which combines a world model with an action policy.
- The financing reality is blunt: “Companies in every direction are actively raising in the primary market, and they go bang-bang-bang-bang-bang straight up.” That is why he calls world models a textbook example of FOMO.
25. World Models and LLMs: Nature and Nurture, Not One Replacing the Other
- 严千行 rejects the replacement narrative. Language is a highly compressed channel for interacting with the world, while sensation, touch, vision, and execution form another channel. “At root, they are simply different channels for interacting with the world”; world models extend capabilities on top of language understanding. Why are there so many routes? “People think about the world differently. Those good at abstract reasoning may gravitate toward mathematics, while those good at concrete imagination may gravitate toward art. A unified model may emerge, or it may not—I cannot predict that today.”
- 刘鹏奇’s internal-sharing framework is that world models resemble innate abilities—capabilities for interacting with the world developed through billions of years of evolution—while language models resemble learned abilities, a second growth curve opened after language acquisition. One is innate and one acquired; they evolve differently while jointly shaping human capability.
26. Why the Breakout Came This Year: Better Video Models and Mature 3D Representations
- Supply-side change No. 1 is the steady improvement of video models. From the classic video of a movie star eating noodles to today’s generated scenes, the earlier output had little to do with physics, while current output is highly realistic. Human video naturally contains physical laws as prior data, and ample open-source video-model supply lets teams train to meaningful intermediate results quickly.
- Supply-side change No. 2 is the evolution of 3D representations: mesh point clouds, NeRF, Gaussian splatting, and then VGGT, which can build a 3D model from a single photograph. “The foundation for generating a 3D world is now in place.”
- But the barriers differ sharply by route. A truly 3D virtual world demands extremely high compute and training costs; “only Tencent has released a version domestically,” helped by its strength in games, while 李飞飞 has released some results. Video-driven world models, by contrast, have drawn in large numbers of embodied-AI and gaming teams. How ready the data is determines how crowded the sector becomes.
27. Where Does Physical Data Come From? 3 Routes and an Order of Commercialization
- World models face a more severe data shortage than LLMs. The industry’s consensus data pyramid has synthetic data at the bottom, ego-centric first-person data in the middle, and real-machine teleoperation at the top. After Nvidia proposed its roadmap, everyone wanted to take first-person data “from 100,000 hours to 1 million or even 10 million hours.”
- 严千行 divides the sources into 3 categories. First is real physical-sensor collection, combining force, tactile, and temperature sensors with video streams. Second is simulation and digital twins. He defends synthetic data: given enough time and precision, CAE companies can produce data with more than 95% accuracy, and “that is valuable too,” unlike the sim-to-real gap criticized in robotics. Third is video-based physical inference and labeling, with AI inferring contact and force states directly from video.
- Economics imposes the constraint. “Everyone agrees that real sensor data is best, but I cannot spend RMB50B–60B collecting one dataset to train a model; that is not realistic.” The ultimate answer is a mix of data types allocated according to ROI, creating a gradient of data quality and cost, as in embodied AI.
- 刘鹏奇 predicts that the first commercial impact will be games and content consumption. Absolute accuracy is not required; “it only needs to fool the human eye,” which pure data-driven approaches can achieve. Embodied manipulation will take longer. 严千行’s hope is simple: “If it really works in games, we will never again have to see bizarre mesh intersections and strange bugs in games.”
28. The Young-Professor Startup Wave: Betting on the Next DeepSeek
- Many world-model founders are PhD students who have not graduated or professors who have only just taken faculty positions. The premise is that investors have accepted the new-lab concept: the question shifts from revenue and deployment timing to “what kind of results can you produce that change the world, and then how do we discuss commercialization?” The professors’ motivation is also real. The best research requires industrial-scale money, compute, and people; universities do not have enough resources, so “they have no choice but to come out,” much like model companies raising capital.
- 严千行 does not deny the absurdity in pricing: “There is definitely an element of madness—with no change in the project, 4 or 5 consecutive rounds can send the valuation up 10x. Rationally, that is clearly unreasonable.” But investors are betting on producing “the next DeepSeek in China—a DeepSeek in a different field.”
- 刘鹏奇 provides the critical counterweight. Whether these academics start companies or not, their goal is research; even in the worst case, failure as a startup may still produce a strong academic result—“that may be one hidden risk inside this bubble.” 严千行 quotes his PhD adviser, who only entered academia after selling a startup: “Being a professor and running a startup are completely different. Research at a startup is research aimed at making the startup succeed, not research aimed at publishing papers.”
29. Has VLA Been Thrown in the Trash? Generative Substitution and the Tennis Analogy
- The surface narrative in embodied AI this year is that “VLA has become universally disliked while world models are everywhere.” 严千行 first explains VLA’s problem: in VLMs, the visual component is small and the language component large, so physical understanding is weak; the system relies on imitation learning from teleoperation data, and generalization remains difficult. Some even argue that VLA is fundamentally overfitting. The new paradigm uses an open-source video-generation model as the base: “use generation instead of understanding.” Like a person imagining what will happen before making a movement, the system follows the imagined outcome; physical constraints are already encoded in the video model’s latent space, raising the generalization ceiling.
- But North America’s latest narrative is correcting the overreaction. After Generalist AI released new results, CEO Peter Florence posted “Beyond VLA and World Models.” The core idea is not to use the newest hammer on every embodied-AI nail, but to replace methodology-driven development with goal-driven development and build a system around zero-shot generalizable objectives.
- 严千行’s tennis analogy makes the complementarity vivid: “VLA sees the ball coming and instinctively returns it. A world model first thinks about how it wants to return it—but thinking before moving is slower than reacting. Elite athletes on a court do not think first and move second; they rely on bodily instinct. In an unfamiliar environment, prediction and imagination help. For familiar movements, you do not need a world model.” 刘鹏奇 adds that some companies use world models as environments for training VLA, while WAM combines prediction and action directly. Every form of combination is being tested.
30. The Humanoid IPO Window: Performance Is Commercialization, Scale Is Still Far Away
- Unitree is moving quickly toward a STAR Market listing, while Leju, Deep Robotics, and others are queuing for Hong Kong listings. 严千行’s assessment is restrained but not pessimistic: performances, research, reception, and exhibitions have already created real rental and sales volume. “You cannot say performances do not count as commercialization; drone shows are an important commercial use case.” Unitree has substantial revenue and profit and looks, in classic terms, like a hardware company. 刘鹏奇’s phrasing is that current sales are mostly emotional value—“but emotional value is still value.” Functional value will take longer.
- Higher-value use cases depend on intelligence. Auto-show guidance, with Luna from an unclear vendor, Figure’s logistics loading and unloading demonstrations, and 智元’s factory livestreams are important early experiments. But the huge difference between this generation of humanoids and traditional industrial robots is that traditional machines are driven by programmed human intelligence. Only when models and bodies combine into intelligent products will expectations for truly large-scale commercialization arrive.
31. The Impossible Triangle and the Lab Mindset: Do Research Before the Inflection Point
- 刘鹏奇’s framework is that embodied deployment faces an impossible triangle: long-horizon complex and delicate tasks × high success rates × cross-scenario generalization. All 3 cannot currently be achieved at once, so deployments must choose 1 or 2 within what customers can accept. 严千行 responds with a classic principle: “The stronger the foundation model, the more effective reinforcement learning becomes.” RL adds little value on a weak base model, which is why the industry is strengthening general-purpose base models today. Once a qualitative threshold is crossed, RL and reasoning can produce emergent results far beyond expectations and widen the entire triangle.
- The contrast is worth recording: Chinese embodied-AI companies have been repeatedly pushed by every constituency to attempt deployment, while overseas companies are closer to a lab model—Physical Intelligence, Generalist AI, and Sunday, among others. 严千行’s position is clear: commercialization depends on whether technical maturity has crossed the commercialization threshold. “Genuinely settling down to do research first is more valuable for the long term.” 刘鹏奇 preserves the disagreement: some companies believe they should pursue a phased deployment path in the middle. “We do not know the answer, but we think that possibility exists.”
32. Dexterous Hands and Tactile Sensing: From “Cool Hands” to Hands That Can Pick Up Tofu
- The mainstream view in 2024 was that hands were unnecessary and grippers were enough. 2 years later, that view has reversed: progress in dexterous hands has made ego-centric human-hand data usable. Mapping a human hand onto a dexterous hand is far better than mapping it onto a 2-finger gripper. “Dexterous hands and embodied intelligence reinforce each other: only good dexterous hands make more data usable.” Hardware has converged around 2 routes—tendon cables and joint motors. The KPIs have shifted from degrees of freedom, payload, and precision in 24 to cost, durability, and maintainability today: “Is it a hand that looks cool, or a hand that can actually be used?”
- Tactile sensors remain within the same route structure as in 24: magnetic, MEMS, and vision-based tactile sensing, with the last using redundant information reduction to extract features and attracting the most attention. 刘鹏奇 proposes a capability test: can a robot infer an object’s 3D shape using only tactile information from its hand? 严千行 immediately supplies the mathematical gloss: “You need surface tactile perception, inferring the envelope surface from a set of normal vectors.”
- The most tangible example comes from a force-torque sensor company backed by Fengrui. A 2-finger servo gripper costing RMB100–200, combined with tactile sensing to close the control loop, can precisely pick up a piece of tofu without crushing it. The timing remains uncertain: everyone knows tactile sensing matters, but the model-side ability to use it effectively will likely come later.
33. Underestimated Areas: AI for Science Expands and China’s Soil for AI Hardware
- 严千行’s first area of interest for the second half is a new version of AI for Science—not an AlphaFold-style point solution, but AI replacing steps in the research pipeline to enable auto research. More aggressively, neural-network operators can replace partial-differential-equation solvers, rewriting the old CAE trade-off of “accurate but slow, fast but inaccurate.” It may now be possible to be both fast and accurate. He also points to the social sciences, which lack efficient mechanisms for simulated experiments: either experiment on an entire society or rely on small samples. Whether AI can accelerate that process is an intriguing question. 刘鹏奇 responds: “Suddenly I feel like the possibilities have opened up a lot.”
- The second area is AI hardware. “It is not that we have seen a specific opportunity and know something must emerge; we have the best soil—a mature hardware supply chain, smart young founders, and strong AI infrastructure. I am very excited to see unexpected surprises evolve at random.”
- 刘鹏奇 connects AI hardware back to world models. Physical-world interaction data, biological data, and EEG data are all extremely scarce, so the clearest short-term value of AI hardware is as an entry point for collecting contextual data on real human activity. Founders must define the product carefully: “Is it merely a data collector, or can it serve as a data hub that integrates the user’s context? At root it is also a consumer product, and people have to want to buy it.”
34. The Essence of the IPO Wave: Ecosystem Positioning, Not P/E
- The backdrop is crowded: Unitree is about to list on the STAR Market; Zhipu and MiniMax have performed strongly since listing early this year; Kimi, StepFun, and DeepSeek have completed large financings; and OpenAI and Anthropic may list this year. 严千行’s characterization is that these companies have capitalized and realized their ecosystem positions after demonstrating them at a particular stage. They need a continuous supply of ammunition to compete and fight a longer war.
- He rejects a P/E framework: these companies are not listing on the basis of earnings. “The bigger question is their long-term ecosystem position.” When internet companies listed, investors were not looking only at who was profitable; they were looking at ecosystem advantages built on long-term cash flows. Ultimately, share prices will be determined by market competition and each company’s final ecosystem ranking.
- The transmission mechanism is what primary-market investors need to remember. Secondary-market feedback operates on a much shorter cycle than primary investing. “If these leading names cannot support market expectations, they will in turn suppress valuation expectations across our primary market.” Once primary and secondary valuations invert, the first half’s valuation stampede may enter “a major cooling-off period”: even secondary-market gains may not hold, so why should primary-market investors pay such high prices or keep acting irrationally?
35. Bubble-Breaking Scenarios and the Second Half’s Single Bet: A Real Data Loop
- On when the bubble will break, 严千行 first defuses the cliché: “Everyone can say the bubble will eventually break, but that is a correct but useless statement.” He then models what happens to capital after the break. Among research companies, those with sustained research-delivery capability and a record of meaningful milestones will continue to raise money, while companies relying on demos and storytelling without results will not. Deployment capital will concentrate in companies with more robust business models and more credible orders; commercial resilience will decide the winners. 刘鹏奇 sees an upside: a broken bubble may help deployment by lowering upstream costs and creating better conditions for adoption. Once companies are public, comparability also improves and helps the market screen for quality.
- The variable 严千行 is betting on for the second half is specific: a genuine data loop. In embodied AI, a company that can continuously roll the entire process—from the data system, pretraining, and post-training through deployment, real-machine data, and reinforcement learning—has a natural advantage over a point solution. The same applies to agents: “skills do not constitute a business model.” Skills can be downloaded freely and are not protected; what can be protected is an agent’s process of making decisions, reviewing itself, and learning on its own.
- 刘鹏奇’s forecast is practical: no technology can achieve AGI directly without passing through a cycle. During the volatility, the ability to deploy determines whether a company can establish a temporary advantage—revenue, financing, and a data loop that carry it to the next takeoff. The episode closes by reusing last year’s line: “We hope to be proven wrong,” and hope the sector’s rapid change gives them more new material to learn from and discuss next year.