Pioneers Insight Method Research Author
Language, Coding, Multimodality: Who Actually Gets a Seat at AI’s Main Table? — 98-Slide Ecstasy Deck, Solo
Back to Episodes

Language, Coding, Multimodality: Who Actually Gets a Seat at AI’s Main Table? — 98-Slide Ecstasy Deck, Solo

Summary

  • Over the past 6 months, AI’s main table has clearly shifted toward coding, and the deciding factor is no longer model capability alone: base model + harness = agent. The model is the engine; the harness adds permissions, skills, tools, memory, feedback and task orchestration, turning AI from “something you ask” into an actor that takes action. As the boundary between coding agents and general-purpose agents blurs, coding’s serviceable market is expanding from software development to almost any digital task.
  • Anthropic has overtaken OpenAI on the hardest revenue metric, shifting the AI-leaderboard fight from ChatGPT’s consumer dominance to Claude’s enterprise assault. The figures cited were roughly $30B in ARR for Anthropic versus $25B for OpenAI; cloud-channel fees and other accounting differences may affect the comparison, but the direction is clear. ChatGPT still has roughly 900M weekly active users and 71% monthly retention. Anthropic’s line is: “Revenue represents everything.”
  • The model race has compressed from a 6-month clock to 2 months or even monthly releases, while Chinese firms’ core advantages are open source and cost. GPT 5.5, Claude 4.7, Grok 4.3, Kimi K2.6, Qwen 3.6 and DeepSeek V4 have arrived in rapid succession. DeepSeek is no longer responsible for “breaking the ceiling”; it is targeting the Pareto frontier at lower cost, then cutting prices to 25% of the previous level.
  • The SaaS selloff is not a normal cyclical pullback; the market is questioning whether the entire business model will be inverted by Service as Software. Software briefly became the S&P’s worst-performing sector, down roughly 20%, with some companies losing three-quarters of their value and Figma approaching a 90% decline. A software project that once took 6 months and $500K can be reduced by low-code to 6 weeks and $50K, then by Web coding to 6 hours and $50. “AI eating software” has therefore entered cash-flow and valuation models.
  • The risk for the Mag 7 and information technology has shifted from revenue growth to capex, free cash flow and payback periods. Three leading cloud companies plus Meta are spending roughly $300B a year; Microsoft fell about one-third in a quarter, while Nvidia’s forward P/E briefly fell to 17x. Yet the S&P hit a new high on April 17 and Nvidia returned to a $5T valuation, reinforcing the view that even violent pullbacks may be reversed very quickly.
  • The compute trade is spreading beyond GPUs into CPUs, storage and optical components, but the physical world is becoming the hard constraint on valuation realization. The server mix in the agent era is described as moving from roughly 1 CPU to 7–8 GPUs during training, and 1:4 during inference, to 1:1 for agents. Intel, AMD and ARM are being re-rated, but manufacturing, packaging, power, permitting and construction delays mean “building 1GW is not that easy.”
  • Private markets are even more euphoric—and more concentrated—than public markets: Q1 2026 investment reached roughly $300B, with the top 5 funds and top 5 deals each taking around three-quarters of the total. OpenAI is valued at more than $800B; one Coatue investor’s 2030 target for Anthropic is $2T; and SpaceX and OpenAI are already among the world’s 15 most valuable companies despite being private. 庄明浩 does not call the top of the bubble, but notes that if the market resembles 1999, “the craziest growth phase” may still lie ahead.
  • The AI euphoria is colliding head-on with employment, energy and social sentiment, shifting the focus to whether economic growth is decoupling from wages and benefits. Daily ad spend on AI short dramas is nearing RMB100M, while multiple U.S. states are discussing data-center restrictions and an “Engels pause” may be repeating. But 庄明浩 rejects compressing a person’s 40 years into “a page of history,” ending with: “Let’s give it the things that are right and cannot go wrong; we’ll do the things where it’s possible to make a few mistakes.”

Deep dive

1. AI’s Narrative Cycle Has Compressed from 3 Months to 30 Days

  • 庄明浩 used to survey the industry once every 3 months: after presenting the November 2025 edition, he produced an annual review in January 2026, then updated it again on March 9 after the Spring Festival. Before 2 months had passed, the PPT had to be rebuilt. “Thirty days in the east, 30 days in the west—not 30 years, 30 days.”

  • The pace is so fast that even the title expires. In early April, the working title was “Can you raise a lobster?”; by the time production was underway, it had become “Is your lobster still alive?” OpenClaw’s rapid transition from fresh phenomenon to “something from the previous era” captures this acceleration.

  • He retains the three-table framework: language corresponds to ChatGPT; coding to the possibility that “AI may write all code within a year”; and multimodality to the idea that “the world model is the womb that nurtures AGI and enables infinite expansion of learning.” This time, he uses technology, industry and capital as the three lenses for deciding who gets the main table.

2. Agents Have Gone from Annual Theme to Industry Default

  • From Qwen’s “From Question to Action” to Tencent’s “AI is moving from chatbot to agent,” 庄明浩 argues that the industry’s biggest 2026 consensus needs no further proof: AI is moving from answering language prompts to executing tasks, and the unit of value is shifting from a single conversation to a completed action.

  • OpenAI’s capability ladder provides a timeline: L1 chatbot, L2 reasoning, L3 agent, L4 researcher, L5 organizer. GPT-3.5 was driven by pretraining; o1, released in September 2024, moved post-training and reinforcement learning to the center; and after DeepSeek’s arrival in late January 2025, L2 reasoning became standard.

  • At L3, the training objective becomes “agentic thinking”: the model must decide when to stop thinking and when to act, select tools and the order in which to call them, extract observations from noisy environments, revise its plan after failure, and remain coherent across multiple rounds and tool calls.

  • 庄明浩 links 林俊旸’s post-departure essay to the Silicon Valley consensus: all of these capabilities ultimately point toward autonomous learning by models. Prompt engineering, context engineering and reinforcement-learning environments remain important, but they are now lower-level components of a larger execution system.

3. The Harness Puts a Powerful Model into a Car

  • Ahead of AI’s analogy is the clearest explanation of the session: an L1 model is an engine that can turn; L2 is an F1 engine. But for an agent to get on the road, it still needs wheels, a driveshaft, a steering wheel, a chassis and tires. That entire structure outside the model is the harness.

  • The formula cited by 庄明浩 is “base model + harness = agent.” A foundation model can reason, generate and understand, but is “static, passive and directionless.” The harness uses structure, direction and constraints to narrow infinite possibilities into finite, purposeful activity, turning AI from an object that is asked questions into an actor that takes action.

  • A usable agent vehicle needs at least permission guardrails, skills packaging, tool calling, a memory system, evaluation and feedback, and task orchestration. It must remember who the user is and what standard the task requires, while continuously looping and self-correcting.

  • This redraws the boundary of infrastructure. Projects serving agents were previously labeled AI infra by the primary market; now context management, memory, continual learning, guardrails, orchestration and evaluation are all being folded into the harness. “Harness is the new infra.”

4. The Agent Ecosystem Is Swallowing the Boundary Between Software Development and General Applications

  • World models are another branch gaining momentum. The industry wants capabilities beyond language that can touch video streams, 3D environments or virtual worlds. The effort can be traced back to 2019 and has more recently expanded to World Labs, Tencent Hunyuan, later versions of Seed 2.0 and related models from Alibaba.

  • OpenClaw is itself a typical harness: an IM application connects to a gateway, which connects to models, multiple agents, memory and tool calls, before splitting into interaction, gateway, agent and execution layers. It does not build the model; it builds the shell and the vehicle.

  • After Claude Code’s code was “accidentally” leaked, outsiders dissected its multilayer architecture and reached the same conclusion: its capabilities come from the combination of the Claude model and a strong harness. 庄明浩 relays the industry joke that “Anthropic has become the real OpenAI,” because OpenAI today is not particularly open.

  • In Q1 2026, almost all the attention on GitHub’s trending projects flowed toward agents: roughly 40% of related projects centered on the agent-model ecosystem. Expanding the sample to the 1,000 most-followed projects, a rough tagging exercise found about 81% related to agents; the recurring keywords were agent, Claude, code, skill and OpenClaw.

5. Coding Has Become AI’s Largest Revenue Pool and Broadest Battlefield

  • Coding has moved from GitHub Copilot-style autocomplete to full generation, code review, debugging and requirements management. The revenue gap is even clearer: AI revenue from coding is larger than that of the next three categories combined. The rest include legal, customer service, healthcare, search and writing, making coding “the absolute main table.”

  • The volume of code committed through Claude Code on GitHub more than tripled from the start of the year to the end of March; over the same period, new websites, iOS apps and GitHub code all increased. App supply, dormant for years, is reaccelerating as development capability is democratized: “Even if someone has already made it, I still want to make my own.”

  • Codex’s latest weekly active-user figure is roughly 4M. Claude has continued to produce an “unstoppable” growth curve around Claude 4 and Claude 4.5. Leading agents can now run autonomously for more than 10 consecutive hours, shifting the unit of software production from person-days toward agent runtime.

  • The larger shift is the convergence of coding agents and general-purpose agents. Claude Code, Codex and Trae can make PPTs and organize information; Manus, Genspark and Coze can also complete tasks by writing code. Users do not care about the implementation path. “As long as it solves the problem, that’s enough.”

6. Anthropic Overtakes on Enterprise Revenue While ChatGPT Holds the Consumer Side

  • The latest ARR figures cited were roughly $25B for OpenAI and $30B for Anthropic. The comparison may differ depending on whether cloud-channel fees are deducted, but Anthropic’s growth slope has been visibly steeper since the second half of 2025, with the lead arriving in Q1 2026 ahead of market expectations.

  • ChatGPT’s classic internet metrics remain exceptionally strong: roughly 900M weekly active users, monthly retention of about 71% and still rising, and leadership in both DAU/MAU and total time spent. Gemini has expanded its share, but its absolute scale has not challenged ChatGPT.

  • Anthropic’s response is that these metrics are “not important.” Its ARR grew roughly 10x from the end of 2025 to the beginning of 2026, or about 16x if March alone is used as the reference. Enterprise coverage is rapidly closing in on OpenAI, while the default preferred model crossed over in January 2026; after that, the two companies diverged.

  • The release cadence shows how intense the trench warfare has become: Claude shipped more than 70 features and products in 50-plus working days, averaging roughly 1.5 per day. “You wake up, open your eyes, and Claude has released another version.” Today’s version kills yesterday’s; tomorrow’s kills today’s.

7. Both Model Companies Are Going All in One

  • OpenAI is gradually merging ChatGPT, Codex and the browser. The Codex App opens directly to a natural-language dialogue box, and the latest version has a browser built in. Three products that once sat side by side internally are converging on a single task entry point.

  • Anthropic is commercializing its first-party harness: it offers both a leading model and the Claude-based vehicle around it, and charges directly at that layer. Third parties can still build shells, but a model company is unlikely to surrender the entire profit pool.

  • 庄明浩 sees 2 organizational paths to the same end state. OpenAI is expanding from a consumer super-portal into coding and the browser; Anthropic is extending from models and enterprise services into execution systems. Both are aiming to own the entire chain from intent to task completion.

8. Model Releases Have Compressed from a 6-Month Generation to Monthly Iteration

  • In the month before the Spring Festival, China’s leading vendors released a concentrated run of updates: ERNIE 5.0 on January 22, Kimi 2.5 on January 27, Step 3.5 on February 2, Zhipu GLM-5 on February 11, MiniMax 2.5 on February 13, Seed 2.0 on February 14 and Qwen 3.5 on February 16. Previously, major versions still arrived roughly every 6 months.

  • By mid-March, MiniMax had already released 2.7, barely more than 1 month after 2.5. April then brought Qwen 3.6 Plus and Claude 5.1; the presentation later refers to “Qwen 5.1” as having launched on April 7, leaving the naming unclear. The list also includes the 35B version of Qwen 3.6, Claude 4.7, Grok 4.3, Qwen 3.6 Max, Kimi K2.6, GPT 5.5 and DeepSeek V4. “From here on, it’s monthly.”

  • 庄明浩 cautions that the released version is often just the previous generation that has completed training. Vendors also have one trained generation waiting to be tested and another in training. That is why more than 10 vendors are packed tightly together at the back of the rankings, while the leading edge moves up almost every day.

  • Sentiment around GPT 5.5 has reversed, and DeepSeek V4 has finally delivered on expectations. In the March edition he wrote, “We’re not waiting for V4.” The change this time is not rhetorical; the release window has become short enough for a single presentation-preparation cycle to span an entire generation.

9. DeepSeek and China’s Open-Source Models Are Targeting the Cost Curve and Ecosystem Control

  • DeepSeek V4’s significance is not that it once again “breaks the ceiling,” but that it achieves an extremely high score at relatively low cost, with its smaller version sitting on the capability-price Pareto frontier. After token pricing was announced, it quickly cut prices to 25% of the previous level, continuing to treat capability diffusion as its central role.

  • On the claim that a model is “too dangerous to release,” 庄明浩 believes there is “certainly some marketing in it.” But he leaves room for the opposite possibility: if a model is powerful enough to discover a software vulnerability that no one had found in 27 years, conducting more safety tests before release may genuinely be necessary.

  • Download volume, model share and capability rankings all show Chinese vendors to be highly competitive in open source, with Qwen and DeepSeek as the leading examples. Hearings in the U.S. Congress described America’s absence from open source as handing the field to China. Open source is therefore no longer merely a software strategy; it has become a theater of U.S.-China AI competition.

10. Sora’s Shutdown Is Not the Endgame; Image 2 Puts Knowledge and Images at the Same Table

  • A16Z’s traffic rankings still show images and video as the largest application category, but 庄明浩 believes DAU rankings are losing value. After Sora shut down around March 28 or 29, the conclusions that retention was weak and “AI TikTok does not work” were probably correct, but they may also be empty truths of the “people eventually die” variety.

  • The more important question is OpenAI’s dynamic trade-off: how its product and organizational structure changes, how finite tokens are allocated among training, research and user service, and when a model company with a particular valuation and competitive position chooses to contract. Sora was shut down to avoid an overextended product sprawl and falling further behind Anthropic.

  • Image 2 then produced a reversal. A single prompt—“generate a screenshot of a Douyin livestream”—can produce a middle-aged woman selling a RMB39.9 Beijing day tour in front of the CCTV building, while accurately handling comments, likes, red packets, viewer counts, rankings and the shopping-cart icon. The model is demonstrating an understanding not merely of images, but of product know-how.

  • Cross-over posters for Genshin Impact, Wuthering Waves, Rock Kingdom and Black Myth: Wukong are even more striking. The model generates the visuals, adds copyright notices for miHoYo, Kuro Games, Tencent and Game Science, and handles capitalization correctly. “Multimodality and language are at one table”; on world knowledge, Image 2 even appears to go a step beyond Nano Banana.

11. Vertical AI Is Monetizing First in High-Value Knowledge Work

  • The time required to go from $0 to $100M in ARR continues to shrink, with legal-AI companies such as Legora and Harvey setting new records. In less than 2 years, Harvey went from a $50M Series A to a valuation of roughly $11B, becoming the classic case of an investment “you should have made at the time, but had no way to make.”

  • CIO surveys show that companies most want to cut IT and SaaS budgets, while security is the budget they are least willing to cut; the released capital continues to flow into AI. Under the substitution framework cited by 庄明浩, the most vulnerable work is already outsourced and relies primarily on information search. Insurance and IT services are first in line.

  • Outsourced work that depends on human judgment looks more like Copilot in the short term, while internally staffed judgment roles remain under observation. The next wave may target internal information work. The conclusion is not that any industry is absolutely safe, but that the issue spans “almost every vertical; it is only a question of sequencing.”

12. SaaS Valuations Have Collapsed Because the Business Model Is Being Rewritten in Reverse

  • As of the end of February 2026, software was the S&P 500’s worst-performing sector, down roughly 20%. Many well-known SaaS companies had lost three-quarters of their value; Figma was down more than 80% and could be approaching a 90% decline. Some companies had fallen to roughly 10x P/E and about 4x forward P/S, even with gross margins still near 70%.

  • ServiceNow’s revenue continued to grow steadily, but its P/E fell from a peak of roughly 132x to 23x, followed by another decline of more than 10 percentage points after earnings. Since Q2 2025, “AI” has been mentioned more often than “earnings” on earnings calls, and the market has begun refusing to pay for traditional growth simply because it is being delivered.

  • That conflicts with software’s long record as the model sector: average revenue growth of roughly 17% versus about 6% for the rest of the S&P, and average gross margins of about 74%, with subscription models providing exceptional predictability. Precisely because the fundamentals remain solid, this selloff is not simply an ordinary earnings recession.

  • 庄明浩’s cost ladder is: 6 months and $500K for traditional development, 6 weeks and $50K with low-code, and 6 hours and $50 with Web coding. “Software as a Service” is therefore reversing into “Service as Software”—the user buys the outcome, while the software itself can be generated on demand.

13. The Mag 7 Is Being Re-rated on Free Cash Flow, Not Revenue

  • The Mag 7 led the market in 2023, saw its advantage narrow in 2024, gave way to AI software and energy in 2025, and fell across the board in Q1 2026. Microsoft dropped roughly one-third, its worst quarter since 1998. 庄明浩 says that aside from Copilot not being good enough, Microsoft did nothing wrong; the market is focused on return on investment.

  • Three cloud companies plus Meta, which is investing “without regard to the cost,” are spending roughly $300B a year in capex. Revenue and operating cash flow can continue to grow, but free cash flow is being consumed by data centers and cannot be recovered in the short term. The market is therefore repricing the entire Mag 7.

  • Nvidia’s revenue continues to grow and both earnings reports beat expectations, yet the stock once traded sideways at $180–$200 for 9 months, with forward P/E briefly reaching only 17x. Microsoft, Amazon, Google and Meta trade at roughly 20–30x, while Walmart and Costco command 40–50x despite expected revenue growth of only 4%–9%.

  • The contrast raises the more fundamental question of what the market is buying—and what it fears. Nvidia’s revenue is expected to grow roughly 70%; a plan cited at the event even points from $200B-plus in 2026 to $1T in 2027. Yet investors still prefer to pay a premium for the certainty of traditional consumption.

14. Once the Tech Premium Vanished, the Pullback Was Erased in a Week

  • In Q1 2026, more than 90% of information-technology companies beat revenue expectations, yet Nvidia, Apple, Qualcomm, Microsoft, Google, Amazon, Meta and Tesla all generally fell. Per-share earnings for the S&P continued to rise, while free cash flow declined visibly, and tech-stock P/Es fell back toward the S&P average.

  • The 2 clearest remaining trades are storage and optics. SanDisk, Western Digital, Seagate and Micron led, while optical-module companies rose in parallel; 中际旭创 in A-shares was driven by the same logic. These are the most direct delivery points for data-center capex.

  • By April 17, however, the S&P 500 had returned to a record high, while Nvidia broke above $200 and returned to roughly a $5T valuation. 庄明浩 cites 王慧文’s warning that humans tend to oversimplify probability and use historical experience to predict the future, then reiterates his own view: even a violent pullback may be reversed in an extremely short period.

15. Data Centers Are Becoming Token Factories, Pulling CPUs Back to Center Stage

  • In 2026, U.S. spending on data centers exceeded spending on offices, and the 2 curves will continue to diverge: “offices for the silicon world are surging; offices for humans are falling.” Over the next 5 years, AI training and inference may account for 60%–70% of data-center demand, with tokens directly corresponding to revenue.

  • The compute controlled by Microsoft, Google, Amazon, Meta and Oracle accounts for roughly 60%–70% of the market. Even treating China as a single entity, its scale is only roughly comparable to Oracle. Rental prices for the older H100 still rose last quarter, showing that supply and demand remain far from balanced.

  • A training data center typically has roughly 1 CPU for every 7–8 GPUs; inference and reinforcement learning move to about 1:4; agent orchestration and task processing move to 1:1. Harnesses require more peripheral compute, making CPUs more important.

  • Intel rose about 20% after earnings and AMD followed with a gain of about 14%. Year to date, Intel has doubled, ARM is up roughly 68% and AMD about 50%. Intel’s share price has moved above its 2000 dot-com-bubble high, with P/E above 100x and near Cisco’s 117x at the height of that mania. 庄明浩 defines it only as a short-term, high-volatility trade for investors with the courage to take it.

16. Physical Constraints Will Cap the Grandest Compute Narratives

  • Cloud providers can raise capex quickly, but semiconductors and data centers remain constrained by advanced manufacturing, packaging, production-line design, government approvals, labor, electricity and generators. “Building 1GW is not that easy.”

  • Data cited at the event showed that more than half of U.S. data-center projects were delayed in 2025. Most of the sites in the Stargate plan showed limited progress in satellite imagery; the original roughly 5GW narrative may translate into only about 0.3GW actually operating after nearly 2 years.

  • Data centers are also facing social pushback. Michigan passed a construction pause running through November 2027, against a backdrop of local electricity prices rising roughly 17.6% in a year. Around 11–12 U.S. states are seriously discussing bans or temporary suspensions.

  • 庄明浩 treats these constraints as costs omitted from the technology story: electricity prices, the environment, communities and quality of life will not automatically submit to model scaling. Protests around AI have already emerged, and Sam Altman’s residence has reportedly been attacked repeatedly with Molotov cocktails and gunfire. The conflict is moving from financial statements into society itself.

17. Private Capital Is Pushing AI Toward Greater Concentration and Closer to Bubble Territory

  • Only 4 months into 2026, the primary-market bar chart already resembles a full year. Q1 investment reached roughly $300B, a quarterly record that even exceeded the liquidity-fueled 2021 period. The top 5 funds raised roughly three-quarters of the capital, and the top 5 projects captured roughly three-quarters of the investment.

  • The private market’s own “Mag 7” is almost entirely AI: SpaceX, OpenAI, xAI, Databricks and Anthropic, plus Stripe, which is also benefiting from AI-driven payment growth. At YC, AI moved from sharing the stage with other categories to “AI and nothing else,” accounting for more than 95% of projects.

  • OpenAI’s valuation of more than $800B corresponds to roughly $25B in revenue. SpaceX and OpenAI are private yet valuable enough to rank among the world’s 15 most valuable companies. One Coatue investor’s target for Anthropic is $2T in 2030, while Anthropic has already approached a $1T private-market valuation; on the same day, Google committed another $40B investment.

  • 庄明浩 declines to say whether the market is at the foothills, halfway up the mountain or at the summit. If 1999 is the comparison, the bubble is indeed closer—but that also means “the craziest growth phase” may not have arrived yet. The primary market’s circular financing network only makes the picture more complicated.

18. The AI Euphoria Ultimately Comes Down to How People Spend Their 40 Years

  • Books about OpenClaw stand upright in bookstores, while books about Doubao can only lie flat; a few months later, both may be displaced by the next product. 庄明浩 interprets bestseller status as “massive anxiety,” because the technology level has risen to the neck and the speed of learning can never keep pace with version updates.

  • Hongguo Short Drama announced RMB500M in support for live-action short dramas, then on the same day eliminated separate rankings for AI and live-action dramas. Six or 7 of the top 10 may already be AI dramas. On April 15, total AIGC spend reached RMB99.6892M and may soon settle above RMB100M; the supposed support program instead exposes live-action content as the weaker side.

  • He uses the idea of an “Engels pause” to describe a possible transition: new technology pushes up GDP while ordinary workers’ wages and benefits remain unchanged for 40 years. But he rejects the God’s-eye view that “a page of history is a lot of people’s lives.” Forty years is not a page of paper; it is each person’s complete and irreplaceable life.

  • The share of AI content at the Ecstasy event fell from roughly 12% to 6%, even lower this time. 庄明浩 sees this not as ignorance but as an emerging attitude. As AI becomes better at doing things correctly and without mistakes, “let’s give it the things that are right and cannot go wrong; we’ll do the things where it’s possible to make a few mistakes.”