Vol.92 Revisiting the Early-March Podcasts Feels Almost Like Another Lifetime — A Crossover with 藏金阁
Summary
OpenClaw’s breakout was not a zero-to-one breakthrough in foundational technology; it was an engineering recombination that brought together “give AI a computer,” an IM channel, an open-source ecosystem, and existing technologies at the same moment. 庄明浩 sees it as assembling existing raw materials “like building with blocks” into a product form that happened to fit the moment; core users may not be surprised, but when an Agent is “lying in your Feishu contact list every day,” nontechnical users can experience it for the first time as a human-like entity that remains in continuous interaction.
The clearest commercial window today is to turn OpenClaw’s still-complex chain of account, API, permission, notification, and local-environment setup into a foolproof product. For a simple task—checking smart-glasses inventory at 9 a.m. every day—庄明浩 had to work through browser, marketplace, Feishu, and scheduling permissions; installing the Windows local deployment took 3 attempts and about 1–2 hours. He recommends that ordinary users start with cloud options such as Kimi and MiniMax: local deployment feels more like “this is mine,” but it also leaves users managing more APIs, permissions, and failures.
Agents make Web3’s old narrative around protocols, identity, payments, security, and machine-to-machine collaboration more visible—and bring exponential growth in compute demand. An Agent can enter a virtual office, bet tokens against another Agent, or even have a colleague’s “lobster” repair its own “lobster”; 庄明浩’s view is that the internet was designed for humans from the foundation layer through the interface layer and may now need a whole new stack for AI.
China’s foundation-model race has converged from separate tables for language, images, and video into a single table spanning coding, Agent, and native multimodality. Kimi 2.5 began supporting multimodality, with Zhipu following; at the time of recording, the market was speculating that DeepSeek V4 would join. Qwen proposed text, image, and video as “three in, three out,” but that unified-model goal conflicts with an organizational structure that splits pretraining, post-training, and modality teams. For vendors, compute is the immediate constraint, while data is more likely to create the long-term gap: “Having it does not mean you can make it work, but without it you definitely cannot.”
Compute growth is a consensus, but that does not mean Nvidia still offers excess returns that the market has not priced in. 庄明浩 says Nvidia traded sideways around $160–180 for about 4–5 months because the market could already estimate Big Tech’s capex, orders, and earnings; what gets chased are surprise sources of upside such as memory, TOTO’s ceramic material for data-center chip substrates, and 味之精, the MSG maker. “Without surprises, there is no big volatility.”
Whether model capability becomes productivity depends on whether users can narrow broad requests into workflows with clear acceptance criteria. AI comic dramas scaled not just because video models improved, but because the industry worked out a 5–6-stage SOP covering script, storyboard, text-to-image, image-to-video, and director sign-off; prompts pull the model’s uncontrolled interpolation from “1 to 10,000” back into a controllable range. 庄明浩’s advice to ordinary people is not to chase the installation, but “don’t treat it like a god”: use AI as a workbench and start by finding tasks you repeat more than 5 times a week.
As generation costs continue to fall, people will need more imagination, taste, and aesthetics—and the ability to make the final choice among a flood of outputs. Works on the level of Black Myth: Wukong need key steps to score 99 points; a model’s 80 is still not enough, but most low-demand use cases are already “up to the neck.” AlphaGo’s match against 李世石 is a reminder that humans may again go through underestimation, doubt, and collapse. 庄明浩 also emphasizes today’s actionable skills of asking, planning, and breaking tasks down, before narrowing the future’s core to imagination, taste, and aesthetics.
Deep dive
1. OpenClaw Won Through Engineering Recombination, Not a Technology Breakthrough
庄明浩 summarizes OpenClaw’s underlying logic as “give AI a computer”(给AI一台电脑): Agents had mostly been given virtual machines; hardware vendors had also tried letting them operate real files and permissions, but no breakthrough ever emerged that ordinary users could feel.
The author reportedly believes this could have been done long ago; after waiting more than half a year and seeing nobody do it, he suggests it may simply have been a blind spot in plain sight. Security, permissions, and product liability made big companies cautious, while a veteran programmer built a prototype around his own needs.
It was not a zero-to-one invention like GPT, but an “engineering permutation and combination”: “Many of the raw materials may already have been OK.” It ultimately hit a new configuration at the moment when technology, user cognition, and interaction experience matured together.
2. IM Access and Open Source Amplified the Difference Users Could Feel
Earlier AI products were standalone tools, and using them required a small ritual; once OpenClaw plugged into Feishu and other IM platforms, it was “lying in your Feishu list every day,” while the familiar chat format made it feel like a person who was always online.
庄明浩 acknowledges that long-term core users of AI coding and Agents may find the innovation limited. The real jolt is to non-core users, who for the first time experience AI as more than a webpage—as an entity with a sense of relationship that can keep executing tasks.
Open source added an ecosystem multiplier: nobody would easily organize a “ChatGPT tips conference,” but anyone can deploy and modify OpenClaw, organize events around it, and layer products on top—hence the rapid emergence of local meetups and spin-off projects.
3. Needing Six or Seven Steps to Check Inventory Is the Productization Window
庄明浩’s first task for OpenClaw was mundane: check at 9 a.m. every day whether a sold-out smart-glasses model had come back in stock, so a coupon gifted by a friend would not expire.
Actually getting it to run required six or seven checks across marketplace login, browser permissions, page targeting, execution frequency, and Feishu bot notifications. That shows how far it remains from foolproof and leaves commercial software a window to package complex capability into a delivered user experience.
He installed it 3 times on an idle Windows machine, spending about 1–2 hours; connecting Feishu took about 20 minutes, and an error omitted from the tutorial cost another half hour. A tutorial draws a straight line from A to B, but a real user can hit a fork at every step.
4. Cloud Is for Trial; Local Deployment Buys Trust Alongside Responsibility
Cloud’s advantages are low cost, more guardrails, and simpler configuration. Kimi and MiniMax can already be installed and connected to Feishu or WeCom through natural-language instructions; at the time, MiniMax’s overseas trial version did not even require payment.
The appeal of local deployment is not absolute security, but the feeling that “that thing isn’t yours” in the cloud: users hesitate to hand over account credentials, browser access, and local file permissions, and only after enough small mistrust accumulates do they want a dedicated machine.
That is why 庄明浩 recommends that ordinary users start online. Local deployment requires paying for model APIs, configuring endpoints, and understanding permissions and failures on their own; absent a clear need, there is no reason to rush into it.
5. Skills Expand Capability—and Supply-Chain and Persona Risk
Skills add capability like giving “an elementary-school student—or a blank sheet of paper” a new set of abilities. Scraping Twitter comments may require an account, access permissions, and a ready-made social-media Skill; it can also be taught step by step through multiple rounds of natural language, so personal setups vary wildly.
Jess asked whether a Xiaohongshu service charging more than RMB500 to install the system, connect Feishu, and download Skills was dangerous. 庄明浩 confirmed that the risk was real: third-party sites can already scan lobsters “running naked on the internet,” and can even see passwords for core accounts.
Even without downloading a third-party Skill from ClawHub, an Agent may independently visit websites the user never expected. Moltbook is OpenClaw’s forum, so users need to check what their Agent has posted; assigning it a personality does not mean knowing what it has done.
After exploding from an idealistic personal project, OpenClaw began updating frequently, with shoring up security mechanisms as a major focus. The problem is that “handing it over means you trust it,” while the Agent does not actually understand the complexity of the network environment.
6. AI Avatars First Run into Boundaries, Not Capability
When discussing the app he had registered with Jess’s invite code, 庄明浩 at one point asked, “What was it called again—Alice?” While using it, the AI left details of his recent emotional spiraling with a real person in his life; that person immediately DM’d him to ask about it, prompting the reaction: “I was furious.”
Traditional software can hard-code a rule saying “do not disclose private chats,” but large language models are difficult to constrain with binary guardrails. More awkwardly, some people may welcome that kind of proactive matchmaking; the boundary itself differs from person to person.
A “digital double” has no single definition: it can be a personality double, a work double, the person the user wants to become, or an entity with an independent “personality and soul.” Each choice changes its memory, recommendations, expression, and product liability.
7. Agents Give Web3’s Old Narrative Tangible Form Again
庄明浩 believes the public chains, identity, payments, security, permissions, and recordable interactions once described by Web3 become “fuller and more concrete” in the Agent era. The difference is that specific behavior is now visible, without being financialized from day one.
An open-source project gave a lobster a visual office: it types while executing tasks and lies down when resting. Later builders put 2 Agents on the same screen, had them fall in love and gamble, and let them bet with tokens.
The stack resembles what is recently being called Web4: open source at the bottom, with rooms, collaboration, authentication, and payments layered above. Chinese projects may reject virtual currencies, but the path of technical evolution remains highly similar to Web3.
8. Machine-to-Machine Collaboration Starts Tearing Down Software’s Walls
郑晓 could not code and wanted to change the base model behind a personal “lobster,” but killed it in the process; in the end, a colleague’s lobster had to repair it. “You think, can this really happen? Yes, it can.”
This Agent wave has “broken through many walls,” enabled by language models that can understand natural language and execute tasks. Boundaries around applications, files, and machines that once had to be managed by people are beginning to be pierced by task intent.
The internet—from protocols to browsers, software, payments, and social networks—was designed for humans. If machines become users, security, insurance, authentication, e-commerce, and interaction models may all need to be redesigned as a new stack.
9. GitHub Goes from an Arcane Tome to a Natural-Language Resource Library
Noncoders once knew GitHub held “countless treasures” but got stuck on Node.js, environment setup, and the command line. Now they can ask an AI coding tool to deploy OpenClaw by following a tutorial.
Likewise, a user can simply say, “Install the open-source PDF-splitting software on GitHub for me,” and let an Agent handle the search, download, and environment configuration without first understanding engineering jargon.
Some of the remaining barrier is fear built up through years of experience. Natural language lowers the entry threshold, but it does not eliminate bugs, environment differences, or permission forks.
10. Models Have Taste Too; Prompts Control the Interpolation Range
Jess’s read is that Gemini is the most nimble, GPT the most rational, and Claude the most “sycophantic.” 庄明浩 says everyone will see it differently, because training teams continually judge whether outputs are good or bad during post-training; taste has already been written into the model.
User memory and feedback continue to shape the model, so the same base model can present different personalities. The claim that “taste matters” applies not only to users but also to the training teams deciding which answers match human preferences.
庄明浩 compares it to the gap between 1 and 2 versus 1 and 10,000: AI’s capability is like interpolating between known points. Give it 1 and 2, and the answer lands roughly at 1.5; give it 1 and 10,000, and the result may drift across a huge range.
Prompts, Skills, and workflows all narrow that range. Many users try OpenClaw for a few days and quit—not because the model is wholly incapable, but because they cannot break their work into stages, outcomes, deliverables, and acceptance criteria.
11. AI Comic Dramas Prove Industry Breakouts Need SOPs, Not One-Off Miracles
AI comic dramas began to gain scale in the second half of last year. Video-model progress was the base layer, but the bigger breakthrough was that the industry worked out a five- or six-stage process covering script, storyboard, text-to-image, image-to-video, and director sign-off.
Once it became clearer who owned each stage, how humans and AI divided the work, what went in upstream, what came out downstream, and how to accept the result, the whole chain could finally “run like an assembly line.”
Jess gave Seedance the same long prompt and still got completely different results. 庄明浩’s diagnosis is that the constraint range remains too wide: even if a model understands the shot, consistency, composition, and execution are not yet fully controllable.
12. Seedance’s Wow Factor Comes from Understanding Storyboard Language
庄明浩 attributes Seedance’s edge to storyboarding: Douyin and Jianying have spent years accumulating a record of how raw footage is cut, retained, and transitioned, and that understanding is harder to replicate than simply generating attractive frames.
His argument is not that “having data guarantees success,” but that a model producing director-like output must understand the editorial relationship between source material and the final cut; ByteDance “already knew this.”
Image and video models will continue to internalize shot language, but external inputs still matter for now. The capability level is rising fast; control precision will determine whether a model moves from an entertaining demo into stable production.
13. Memory Portability Exposes a Vulnerable Side of ChatGPT’s Moat
After Claude introduced a feature framed as “empty ChatGPT in 60 seconds,” Jess wondered how memory could be moved out so easily. 庄明浩’s answer: if a model can understand a user through long-term conversations, then some form of summarizable memory must exist. In theory other models can import it too, but how usable it is afterward will still differ.
Over the past year-plus, OpenAI has made a product bet on ChatGPT remembering users’ preferences, personality, and requirements through frequent conversations, creating retention and differentiation. Once a user firmly decides to leave, a migration tool strikes directly at that moat.
庄明浩 linked Claude’s rise to No. 1 on the U.S. AI rankings that week to the Pentagon-Anthropic conflict and OpenAI taking over the related work. It was his on-the-ground judgment about why users migrated, not a claim that model capabilities had suddenly reversed.
14. The Lunar New Year Traffic War Did Not Rewrite the Ranking; the Model Race Has Converged into One Table
The essence of China’s Lunar New Year battle was still Big Tech fighting for users: milk-tea giveaways and order placements created short-lived peaks, but 庄明浩’s summary was, “First stayed first, second stayed second, third stayed third.” The shape of the curves barely changed.
It was a prisoner’s dilemma: everyone knew subsidies might not create lasting value, but once others did it, no one could sit out. The traffic playbook remained classic internet, and the technology ranking was not reshuffled.
China’s pure-play model companies have shrunk from the “Six Little Dragons” to roughly 4–5. Baichuan and 01.AI have reduced investment in the general-model main battlefield; the remaining players, together with DeepSeek, continue chasing the broadest and most frontier capabilities.
In 2025, one could still say Zhipu was betting on coding, Kimi on Agents, and MiniMax on multimodality. By 2026, “all the tables have become one table”: coding, language, Agents, image, voice, and video now compete in the same arena.
15. Native Multimodality Requires Models, Data, and Organizations to Fuse at Once
Having a model generate an infographic for an article on “the 2028 doomsday theory” actually requires searching for the article, understanding its text, summarizing its structure, and outputting an image. This is no longer an image model alone, but an integrated task spanning language understanding and visual generation.
Kimi 2.5 was the first to support multimodality, and Zhipu began following. At the time of recording, DeepSeek V4 had not yet been released, but the market speculated that it would support multimodality alongside Agent capability. MiniMax planned to further consolidate its 2.5 language model with Agent, video, and voice.
Qwen proposed “three in, three out”: text, images, and video can all be inputs and outputs. But that unified training objective conflicts with Alibaba Cloud’s organizational split between pretraining, post-training, and language, image, and video teams; 庄明浩 believes the conflict between organizational structure and the unified goal is tied to the issue at hand.
The Nano Banana Greece check-in example shows the value of fusion: give the model only a user’s photo and a place name, and it knows what the location looks like. Older models would still require the user to provide a separate local reference image.
16. AGI Has No Single Finish Line, and US-China Competition Is Not the Same Balance Sheet
Asked when AGI will arrive, 庄明浩 answered, “I don’t know; nobody knows.” Even the standard is unsettled. Asking a model to see only material published before 1911 and derive relativity on its own looks measurable, but in practice it is impossible to separate the contributions of data, guidance, and reasoning.
The U.S. still has the world’s top engineers, financial system, energy, and innovation machinery to support massive investment. 庄明浩’s counterpoint is that the U.S. may “have only this left”; if it loses the AI race, it will have fewer industrial pillars to fall back on.
China has more than software AI: manufacturing, the EV supply chain, embodied intelligence, commercial space, and other themes mean its capital market is not solely a bet on foundation models. In China’s policy and market debate, software AI is instead “relatively farther back” on the priority list.
17. China’s Tech Financing Is Forming a Policy-and-Capital-Market SOP
庄明浩 describes the path taken by domestic-chip, embodied-intelligence, and commercial-space companies as a form of “Chinese-style financial-market support for technological innovation”: the businesses are still early, yet quickly receive support from policy, capital, and market narratives.
That explains why investors focus so closely on “15th Five-Year Plan” keywords, and why A-shares can back multiple technology themes simultaneously. He does not judge the mechanism as inherently good or bad; he emphasizes that it can already be replicated repeatedly.
National security is not a binary choice. AI-model companies raise dollars, and overseas institutions will study Chinese AI; Zhipu’s background might suit A-shares, yet it ultimately listed in Hong Kong, illustrating the gray zone for software AI between regulation and capital.
18. “Wrappers” Have a Window of Value; Secondary-Market Prices Can Feed Back into Primary Valuations
庄明浩 rejects treating “wrapping” as categorically worthless: a narrow API wrapper is easy to copy, but turning accounts, permissions, privacy, workflows, and UI into a usable product can still command a window of opportunity.
In the broad sense, ChatGPT itself is a “wrapper” around GPT-3.5: a model cannot serve the mass market directly, and the outer layer of interaction and experience is not optional. The real question is not whether there is a wrapper, but how much hard-to-replicate product capability has accumulated inside it.
He says directly that MiniMax and Zhipu share prices are “notional,” because their moves do not directly add to company cash flow. But rising prices lift Kimi’s valuation, while Kimi’s financing brings in “real cash,” so it has no need to rush an IPO.
19. Compute Is the Immediate Bottleneck; Data Is More Likely to Open the Long-Term Gap
庄明浩 uses the familiar algorithm-compute-data framework: since GPT was released, foundational algorithms have seen “almost no major innovation,” mostly marginal optimization. Compute is constrained by chips, power, land, construction cycles, and U.S.-China restrictions.
OpenClaw’s continuous execution pushes token consumption up exponentially. After Zhipu released GLM-5, the OpenClaw integration wave arrived and VIP access began to be rationed—not because Zhipu did not want to sell, but because it “could not serve that many users.” Some AI public courses likewise had to refund customers because their services could not handle the volume.
Even several hundred thousand users generating videos on Seedance would be enough to make ByteDance impose queues of several hours. Against Douyin’s roughly 800M users, the demand is still very early, yet it is already stressing large companies that are trying to buy as much compute as they can.
Data is the more durable source of differentiation: advantages built on data in the broadest sense are hard to catch. Good data does not guarantee a good model, but “without it, you definitely cannot.”
20. The Market Trades Surprises; Humans Ultimately Bet on Taste and the “Divine Move”
Jess asked how ordinary people should bet on AI, and 庄明浩 offered no advice. Nvidia stayed in the $160–180 range for about 4–5 months—not because compute demand had disappeared, but because Big Tech’s capex, orders, and profit growth were already baked into expectations.
The market instead chased areas few had fully anticipated: memory; TOTO ceramics as an important material for AI data-center chip substrates; and 味之精, the MSG maker. “Clear compute growth” and “a stock still having a surprise” are two different things.
For ordinary users, he recommends “don’t treat it like a god”; treat it as a workbench instead. Find a work, fitness, diet, or study task repeated more than 5 times a week and let AI optimize one segment—there is no need to create new anxiety around installing a lobster.
The AlphaGo documentary offers the deeper mirror: 樊麾 and 李世石 both went through underestimating the opponent, doubt, and shattered conviction; in the fourth game, 李世石 played the intuitive “divine move”(神之一手)that turned the game around, but it also became humanity’s last single-game win over AlphaGo. 庄明浩 had earlier identified asking, planning, and task decomposition as the skills that work today; in his final answer on future capabilities, he distilled the core to imagination, taste, and aesthetics—when AI instantly gives you 5,000 options, you can still say, “This is the one.”