Vol.65 AI’s New Era: Is Google Back?
Summary
Google has returned to the AI first tier; the key takeaway from this I/O is not a one-off leaderboard upset, but that it has finally begun rebuilding core businesses such as Search, Chrome, Gmail, and Android around AI. The episode offers different views on the “real inflection point”: 庄明浩 emphasizes the organizational merger, while 挨踢牛魔王 emphasizes the early release of Gemini 2.0. But this year’s conference showed models, products, and ecosystem distribution beginning to form a closed loop, easing the long-standing market discount on Google as “technically strong but unable to execute in applications.”
Veo 3 has moved the industry standard for AI video from “generating images” to native simultaneous generation of audio and video, putting it temporarily “half a step” ahead of 可灵. 朋克周 even mistook a video of a kangaroo carrying a boarding pass onto a plane for a real prank; by his estimate, generating one through the API costs about RMB 100. It is not yet mainstream, but it can already handle dialogue, ambient sound, eye contact, and scene matching at the same time. “I’ve been following AI news for 2 or 3 years, and even I, an old bagholder, didn’t realize it was AI-generated.”
Gemini 2.5 Pro is the foundation of Google’s entire AI product stack, with strong writing, coding, and reasoning capabilities; Gemini Diffusion points toward a low-latency direction of roughly 1,400–2,000 tokens per second. 挨踢牛魔王 considers the latter of limited practical use for now, but if Search and interaction can shift from character-by-character output to “instant delivery,” product boundaries will be redrawn. 朋克周 tested the updated 2.5 and generated a playable game with a complete UI in 2 minutes.
Google’s most consequential move for investors is to cannibalize the search cash machine itself, rather than build a segregated AI premium tier. The episode uses Bing as a reference point: even with roughly 3% of global search share, it can still generate about $12B in annual revenue, so losing even 1 or 2 points would be material for Google. By launching AI Mode in core US Search, Google is making good on the principle: “If anyone is going to disrupt me, it has to be me.”
DeepSeek and Kimi’s exploration of reasoning-model training rapidly eroded the O-series advantage that OpenAI might otherwise have retained for 1 or 2 years, indirectly lifting the capability curve for Google and the industry as a whole. 庄明浩 summarizes 2025’s progress as “2 consecutive kicks”: reasoning became standard in Q1, then the industry rediscovered room in pre-training in Q2, on top of post-training and reinforcement learning. The market therefore expects V4 and R2, following V3-based R1-0528, to each deliver another leg up in succession—but that remains speculation and anticipation.
The safe zone for startups is not a feature that models can never cover, but a combination of stronger models, better delivery, codifiable industry know-how, proprietary scenario data, or a meaningful cost advantage. 挨踢牛魔王 warns founders not to “stand in front of the frontier models,” because a single upgrade can swallow months of work. 庄明浩 adds: “The model is the application, but applications are not limited to models.” 美图, HeyGen, and vertical voice services share the same path: selling a concrete result rather than generic tokens.
In 2025, the main battlefield has shifted from competing purely on models to Agents, coding, multimodality, browsers, and AI hardware, while the capital threshold is rising sharply. The market benchmarks cited on the show suggest that vertical companies reaching $100M ARR may command 30–50x valuations; Cursor is also rumored to have raised $900M at a $9.9B valuation on roughly $500M ARR. The true long-term variables are OpenAI’s L4 “innovation” and L5 “organization”: technology only becomes productivity when it is combined with organizational and business-model reinvention.
Deep dive
1. Google’s comeback starts with its rehabilitation from industry punchline to clear first-tier contender
潘乱 starts by laying out the full list of scars: Bard’s debut botched a basic factual question, wiping roughly $700B off Google’s market cap overnight; later demos were accused of being doctored, while generated content became embroiled in controversies over racial bias, poisonous mushrooms, and more. “Anyway, the past 2 years have been one scandal after another.”
The more dramatic episode came when Google planned to showcase Project O, which could see, hear, and converse in real time, only for OpenAI to release GPT-4o the day before. Combined with challengers such as New Bing, Perplexity, and Dia, the narratives that “an elephant can’t turn around,” that a hero is past his prime, and that large companies suffer from bureaucracy all settled on Google.
After this year’s I/O, the guests did not declare that Google had overtaken everyone across the board, but the consensus was clear: “Whether it has overtaken anyone is another question”—it can no longer be excluded from the global first tier. In-house TPU, Gemini 2.5 Pro, Veo 3, Search, Chrome, and glasses together amounted to a full-stack showcase.
2. Veo 3 turns “indistinguishable from reality” from a marketing slogan into a real risk
What left the deepest impression on 朋克周 was a video of a woman boarding a plane with a kangaroo. The kangaroo was even carrying a boarding pass in its mouth. When he sent it to his girlfriend, he assumed it was a funny prank by a foreign creator; only after she said, “Look at its hands,” did he realize it might be AI. “I just froze.”
The signal was not only image quality, but the native realism of the entire clip: scene sounds, speech, eye contact, and movement all connected naturally, with very little of the usual AI feel. 朋克周 believes that older people, children, and professionals alike may struggle to distinguish such content consistently—enough to “cause a certain degree of chaos.”
Cost remains a hard constraint. By their estimate, generating a clip through the API costs roughly RMB 100, still far from low-cost production at scale. 挨踢牛魔王 therefore believes Veo 3 may currently be “half a step” ahead of 可灵, rather than having already ended the competition.
3. Native audio-video generation resets the passing grade for every video model
庄明浩 believes Veo 3 solves the core break in the previous workflow: video models produced the image, while sound had to be generated, dubbed, and aligned separately. It turns “video is video and speech is speech” into simultaneous generation, forcing every competitor to accept a new benchmark.
挨踢牛魔王 summarizes the technical path as V2A, a line of research DeepMind has pursued for years: first extract scene information from the video, then use diffusion to generate matching sound. Forest ambience, rain, dialogue, and movement no longer depend entirely on separate post-production, but may be matched to one another during generation.
For production teams, the savings go beyond dubbing costs. In the past, a shot might have to be “rolled” hundreds of times before one was selected, edited, and dubbed. Native audio-video generation gives small companies and independent directors a chance to deliver professional footage that once required expensive equipment and large teams.
4. Gemini 2.5 Pro is the foundational asset behind the entire launch
挨踢牛魔王 selected Gemini 2.5 Pro as the strongest release: it is powerful across writing, coding, and reasoning, and the version released on June 5 had already climbed to the top of the leaderboards. Veo 3 is more spectacular, but Gemini is the foundation shared by Search, Agents, and a wide range of applications.
朋克周 offered more tangible evidence than leaderboard scores. Using the version 2.5 update released that day for coding, he generated a game with a decent UI that was genuinely playable in 2 minutes; the team ended up playing it continuously for nearly half an hour. I/O was no longer just a display of research capability, but began to touch the production speed of usable products.
庄明浩 also noted that Google’s model performance had been gradually catching up with the leaders since 2024. The market had previously refused to give it credit because technical progress had not yet translated into the experience of its main products. The change at this year’s I/O was that model scores and product deployment finally appeared together.
5. Gemini Diffusion bets on an interaction model in which waiting disappears
挨踢牛魔王 considers Gemini Diffusion the most technically novel experiment, though not necessarily the most useful today. It brings the denoising approach commonly used for images into language generation, more like laying down a draft and refining it as a whole than writing forward one token at a time with a traditional autoregressive model.
The episode puts its speed at roughly 1,400–2,000 tokens per second, close to “instant output.” The real product value is not the benchmark score, but that Search, coding, and real-time interaction would no longer have to unfold as slowly as typing; after the user clicks, a nearly complete response could arrive immediately.
This is also seen as a sign that Google’s capacity for innovation is returning. Even if productization is still some way off, Google is at least willing to explore model architectures different from the mainstream autoregressive route rather than merely bolting old technology together into a rushed response like Bard.
6. Google is using its ecosystem to beat standalone products, not building another isolated product suite
朋克周 put it most sharply: “Use my ecosystem to beat your standalone product; use my membership to beat your model.” ChatGPT and Claude remain primarily independent conversational products, while Google can embed AI into the existing user networks of Gmail, Android, Chrome, cloud, Search, and YouTube.
A minimal workflow already shows the difference: use Gemini to generate a prompt, Veo 3 to create the video, Flow to edit it, and then publish it to YouTube with one click. Users do not have to shuttle assets between multiple AI products; the ecosystem itself becomes an advantage in both experience and distribution.
That does not mean every integration is immediately useful. 朋克周 said bluntly that AI in Gmail can sometimes be “slower to edit than writing it yourself”; Google also has so many products that the offering can be dizzying, while Flow was not yet working properly at the time. The advantage is a complete chain, not every node already being polished.
潘乱 described the business model as “88VIP,” and a guest added: “Large volume, fully stocked.” Users may eventually buy not a single model membership but an ecosystem assistant covering work, content, shopping, and daily life—although the price is still high for now.
7. AI-izing Search is a real act of self-cannibalization
潘乱’s economic framing is this: before AI emerged, Google held at least roughly 80% of the global search market; Microsoft’s Bing, with only around 3% share, could still generate approximately $12B in annual revenue. For Google, losing even 1 or 2 points would be enough to create a huge business elsewhere.
This is the core question that has concerned public-market investors over the past 2 years. If Google does not follow AI, Perplexity and others may take share; if it does, answer costs rise and the advertising click path shortens, potentially cannibalizing its richest profit pool. Search advertising may account for an estimated 80%–90% of Google’s business, according to 挨踢牛魔王, making entrenched interests even harder to bypass.
The key signal from this year’s I/O was Google’s willingness to launch AI Mode inside US Search and place AI in the most important search box. 挨踢牛魔王 summarized it as: “If anyone is going to disrupt me, it has to be me.”
8. Search is shifting from an information connector to a results provider
The change is not just a new interface, 潘乱 argues. Traditional Search handled navigation, connection, and traffic distribution; AI now gives a direct answer and may eventually select products, book hotels, and reserve flights. Whether intermediary websites and apps will still receive the same traffic and commercial value is now an open question.
MCP prompted a similar vision: applications may need to shift from “being browsed by people” to “being called by Agents,” with developers providing structured services to models. For startups, value may shift from competing for time spent on a user-facing page to becoming a reliable execution node in an answer or action chain.
For Google, the significance of AI Mode is not merely defensive against Perplexity. It puts the relationship among Search, models, advertising, and shopping back on the table, showing that the company is willing to rebuild its core business with AI rather than isolate innovation in a “premium tier” and protect the old peak.
9. The real inflection point was organizational integration, not a breakout demo
庄明浩 places the inflection point around April or May 2023, when Google Brain, DeepMind, and other AI teams were consolidated under the leadership of Demis Hassabis’s group. Brain was strong in Transformer, TensorFlow, and academic research, but in his assessment it was “good at writing papers, not so good at turning them into products.”
DeepMind had accumulated years of work in multimodality and diffusion, and after the merger it also gained support from the former Brain team in areas such as neural-network optimization. What this I/O showed was not a last-minute sprint, but the staged result of technical, product, and ecosystem capabilities gradually converging after departmental walls came down.
挨踢牛魔王 also mentioned that the semi-retired Sergey Brin had returned to the company to write code and called on the team to work at least 60 hours a week. He believes this gave the broader team a significant morale boost.
10. Gemini 2.0 Flash had already revealed early signs of the counterattack
挨踢牛魔王 places an earlier product inflection point in Gemini 2.0 Flash. It could already edit images, change people’s styles and clothing, and combine people with products. The fact that Google managed to release ahead of the field in image editing suggested a development cadence very different from the Bard era.
OpenAI then overshadowed it with the stronger results of GPT-4o’s image model, making Google look “embarrassing” once again. But 牛魔王 believes the reversal actually began there: Google had regained the ability to release first, rather than merely waiting for competitors to define the direction and then following.
By this I/O, OpenAI’s main simultaneous release was Codex, and it could no longer completely overwhelm Google’s launch as it had in the past. The episode did not use this to declare a winner, but viewed it as a sign that the two sides had shifted from “one-sided interception” to competition at the same level.
11. DeepSeek and Kimi rapidly flattened the moat around reasoning models
挨踢牛魔王 explains the path through “reading and doing exercises.” Pre-training is like reading a vast number of books; reinforcement learning is like doing exercises, training the model to reason on specific tasks. OpenAI’s O series had been thought capable of holding back competitors for 1 or 2 years.
His view is that the Chinese teams behind Kimi and DeepSeek independently found a pure reinforcement-learning route without falling into the Monte Carlo tree approach. Kimi did not open-source its work and drew less outside attention; DeepSeek published its papers and training details, amplifying the impact globally.
牛魔王 interprets OpenAI’s public discussion of Monte Carlo trees as a form of misdirection. That is his inference, not internal evidence supplied by the episode. The core factual chain remains that the 2 teams found a different path from public clues, and DeepSeek’s openness allowed other companies to reproduce and improve it faster.
12. Model progress has become a two-legged engine of pre-training and post-training
庄明浩 uses OpenAI’s L1–L5 framework to map the progression: L1 is the Chatbot, L2 is reasoning, and L3 is the Agent. From the appearance of o1 in September 2024 through Q1 2025, the primary task for leading model companies worldwide was to reproduce reasoning capability.
The industry briefly accepted that the pre-training ceiling had been reached and shifted toward post-training and reinforcement learning. By Q2 2025, it had discovered that pre-training still had room to improve. 庄明浩 calls this “kicking the ball forward”: pre-training lifts one leg, post-training lifts the other, and capability rises in alternating steps.
Market optimism around DeepSeek R2 also comes from this logic. R1-0528 is still based on V3; if V4 first improves the base model and R2 is then trained on top of it, the system could jump twice in succession. 庄明浩 emphasizes that this is a projection and an expectation, not a confirmed release date or outcome.
挨踢牛魔王 therefore believes DeepSeek made “a certain contribution” to Google’s progress at this I/O as well. Open methods helped the industry avoid wrong turns, while Google owns data assets such as Search and YouTube. But this is an inference based on industry diffusion, not a result of direct cooperation between the companies.
13. NotebookLM proves Google can also produce product insights with viral power
NotebookLM did not go viral because knowledge-base Q&A was inherently unique. Its breakthrough was turning dry papers or source materials into a natural podcast conversation between a man and a woman. Users can even interrupt with “I don’t understand this part,” prompting the hosts to stop and explain again, creating the feeling of being accompanied while learning.
朋克周 notes that products such as 扣子空间 can now also turn topics or articles into podcasts, while voice models are rapidly improving in intonation and conversational quality. But NotebookLM was the first to package the capability into a clear, understandable, shareable feature—a particularly meaningful achievement for Google, historically known for strong technology and weak products.
潘乱 gives an example of knowledge aggregation: a friend put 56 interviews with AI-coding founders into NotebookLM and shared the notebook and Q&A link. By continuously asking questions based on the material, users could go down a “rabbit hole” and quickly acquire context close to that of an industry insider.
14. The main line for 2025 has shifted from model demos to Agents and productization
庄明浩’s map is straightforward: Agents are the decisive main battlefield, coding and multimodality are secondary lines, while hardware and social products are smaller areas of exploration. At this stage, the results in consumer entertainment and social products are “not very good.”
潘乱’s commercial categorization is even more direct: coding is currently the clearest proven money-making lane; AIGC is the older main line with the densest participation; Agents are the most crowded direction for startups today; product marketing and Try On are heating up again as models mature.
朋克周 observes that many companies around him have abandoned training their own large language, voice, and even other foundation models, shifting toward applications and hardware. The model layer is increasingly a capital-intensive battlefield for a small number of large companies, while product experience has clearly become an independent variable in 2025.
15. AI coding has become the foundational entry point no major company can afford to miss
朋克周 believes a major company can survive without a browser entry point, but not having its own coding capability is more serious because code is the foundation of the internet and software. After ByteDance advanced Trae, Alibaba also strengthened 通义灵码 and its IDE products; the competition is no longer merely about selling model APIs.
The vulnerability of Agent startups can be seen in a rumor mentioned by 牛魔王. Against the backdrop of talk that OpenAI might acquire a company, a coding tool dependent on the Claude API allegedly faced the awkward prospect of having its supply cut off. The episode did not confirm all the details of the rumor, but it illustrates how reliance on a single model can change the fate of an application company.
朋克周 believes that unless a company has already built market recognition like Manus, or established a brand in a vertical such as design or film, ordinary Agent capabilities will most likely be absorbed by foundation models. Strategic partnerships and distribution therefore matter no less than the feature itself.
The other side is a cliff-like fall in development barriers. A team used Gemini and YouWare for a small hackathon and generated a playable program in 2 minutes. The “best of times” is that anyone can build; the “worst of times” is that everyone depends on the same group of models, accelerating both replication and competition.
16. Monetization moves down the stack from compute to models to applications; timing matters more than imagination
挨踢牛魔王 describes the order of monetization as arriving in waves: first came NVIDIA and the other card sellers, then leading model companies charging by the token, and only afterward AI coding. Once models become better at taking action, Agents will have a more stable basis for paid use.
Entering early does not necessarily mean leading. Early founders once used Stable Diffusion 1.5 for e-commerce clothing try-on, investing millions, tens of millions, or even more than RMB 100M, only to find that the capability could not meet customer requirements. Now that Google and 可灵’s Try On products are gradually maturing, the same demand is beginning to become deliverable.
潘乱 got a market benchmark from friends in cross-border e-commerce: using off-the-shelf consumer models for clothing changes can save roughly 90% of the cost versus certain enterprise-service solutions. This may indicate that capabilities have matured, or it may mean early service providers were charging high prices and that margins will be compressed by general-purpose models.
17. Flow and Veo 3 open up a new production-budget structure, not just more special effects
朋克周 believes that if Flow can make Veo 3 more controllable, paying customers will expand from AI-video enthusiasts to marketing, film, and post-production teams. 可灵 has already been elevated by Kuaishou into an independent first-level business unit and developed strong profitability, suggesting that generated video is moving from demo to business.
潘乱 mentions a short-drama team looking for subjects that are difficult for humans to shoot but well suited to AI: 《灵蛇》, 《山海经》, mecha, and even the snake people and rat people in 燕垒生’s 《天行健》. But a roughly 100-minute short drama with a total budget of RMB 2M–3M still cannot absorb current technology costs. Technical feasibility does not equal economic feasibility.
Model characteristics also shape content in reverse. 可灵 is good at fantasy and mecha, while Veo 3’s strength is ordinary scenes such as a rich person cooking at home; when 朋克周 used it to make science-fiction imagery similar to 《流浪地球 3》, it still distorted. Creators are looking for subjects again based on what each model can actually do.
18. AI turns Google Glass from a failed entry point into a computing platform worth reassessing
挨踢牛魔王 believes the old Google Glass failed because it had too many weaknesses and too little practical use. Multimodal AI fills in perception, knowledge, and reasoning. A user could look at a tree and point at it in the air, and the glasses might understand the object and identify the species without requiring the user to take out a phone and scan it.
More practical scenarios include automatically organizing meeting notes and reminding someone, in a social setting, who another person is. 牛魔王 notes that facial recognition may already be more reliable than human memory; if the glasses have recorded someone before, they could reduce the awkwardness of “they know you, but you don’t recognize them.”
朋克周 takes a broader view: AI is not just rescuing Google Glass, but also bringing a number of previously painful VR and AR companies “back from the dead.” Once speech recognition, visual understanding, and natural interaction are added, glasses, watches, wristbands, cameras, smart homes, and AI toys can all have their uses redefined.
潘乱 pushes the question further: if Gemini leads in multimodality, will Google’s glasses be the best, and could Pixel then surpass the iPhone? The guests did not accept that deterministic chain. They only agreed that glasses “could” replace phones, while computers would be harder; a good model does not automatically produce the best device.
19. Real-time translation shows the clearest immediate value of multimodal hardware
挨踢牛魔王 imagines 4 guests speaking Chinese, English, Spanish, and Hindi, while each person hears their own native language. The system would preserve the original speaker’s voice, tone, and emotion, rather than merely supplying subtitles or mechanical dubbing.
Google has shown a similar conference scenario, but 牛魔王 acknowledges that the product still needs work. His view is that the Transformer originally grew out of translation, and AI may now reverse the process and “gradually solve translation in its entirety.” International conferences and livestreams would benefit first.
20. The browser, a classic product, is once again contested strategic territory for AI
庄明浩 lists the participants as Perplexity, OpenAI—according to rumors—and, in China, 夸克 and Tencent Browser, which is increasing its investment again. A classic product once treated almost as infrastructure is regaining internal resources and strategic influence because it can carry accounts, synchronization, response speed, security, and privacy.
Chrome is also facing an antitrust case in the US, with the outcome still unresolved. 庄明浩 does not expect a sudden “one-shot transaction” to complete a breakup, but acknowledges that long-term uncertainty could affect how Google fights for the entry point.
For browser-extension startups, native AI integration in Chrome would directly compress their room to operate. Account synchronization, response speed, security, and privacy are not capabilities that small teams can easily replicate. But 庄明浩 leaves one opening: “The model is the application, but applications are not limited to models.”
21. AI’s product value lies in inferring hidden parameters, not adding a chat box to an old interface
潘乱 uses e-commerce search to expose a basic weakness. Searching for an address inside Taobao, JD.com, or Meituan could return an atlas for sale; searching for “invoice” could produce a pile of products rather than an entry point for issuing an invoice. User intent is obvious, yet traditional products still depend on rigid keywords.
朋克周 explains that complex software may have 300 parameters behind the scenes. The product displays only 3; the other 297 have not disappeared, but have simply been assigned average defaults. Edge-case users will naturally find the product hard to use.
AI’s potential is to combine the search term, history, and current context to infer those 297 hidden parameters and handle individual differences without adding more form fields. Interaction could therefore shift from “the user learns the software” to “the software understands what the user is trying to accomplish.”
潘乱’s irony is worth preserving: future systems might first “think deeply” about whether the user wants to issue an invoice or buy one, and only then finally display the correct interface. AI can compensate for technical constraints, but basic product experience and intent design still matter.
22. The most dangerous place for a founder is directly in front of the model’s advance
挨踢牛魔王’s warning is blunt: “Do not stand in front of the frontier models. They will crush you.” A startup may spend months working overtime to build a complex workflow for clothing try-on, only for the next foundation-model release to achieve the same result with one prompt, wiping out its equipment, rent, and R&D investment at once.
His proposed defense is to find industry know-how, preferably expressible in code or a clear rule. Using E=mc² as an analogy, a general-purpose model is a universal curve-fitting machine approaching the underlying rule, while an exact formula is more stable and can run on a CPU at a cost far below continually calling a GPU model.
潘乱 then asks whether such know-how will also be absorbed by knowledge bases and models. Using R1 and Kimi as examples, he says neither had access to the complete papers or literature, yet both discovered from clues in OpenAI presentations that reinforcement learning could solve the reasoning-model problem. In this sense, once knowledge bases, models, and Agents are combined, industry experience may also be aggregated further.
牛魔王 ultimately reduces the principle to this: “The stronger the model, the stronger the thing you build,” rather than a model upgrade making the product disappear. If AGI truly covers an entire domain the way AlphaGo swept through Go, he acknowledges that existing defenses could still fail.
23. New applications will not simply be old internet products with an AI layer
庄明浩 uses “独响” to illustrate an extreme product hypothesis: users write posts or diaries, but there are no real people on the platform—only Agents commenting and responding. Users get the feeling of being heard without having to endure disagreement or rejection from actual people. It is unproven, but represents design aimed at future psychological states.
YouWare bets on another change. YouTube lets people share videos; once AI coding lowers the barrier to making applications far enough, people may also share their own games and programs. The smallest unit of a content platform may expand from posts and videos to runnable software.
潘乱 therefore believes it is difficult to conclude that any direction lies entirely outside the intelligence theme. Once knowledge bases, coding, and Agents converge, the only scarce things may be Blade Runner-style “handmade noodles” and “real sheep.” But that is a vision of a trend, not a market outcome that has already occurred.
24. The main lane remains open to startups, but exit paths and capital scales have changed
庄明浩 observes that AI deals in 2024 were often “hollowing-out acquisitions,” essentially talent grabs. In 2025, acquisitions paying substantial prices for companies and products have begun to increase visibly.
Even in the main lane contested by foundation-model companies, startups can still be acquired through execution, product quality, and first-mover scale. Large companies face their own headcount and prioritization constraints and may not build every layer from scratch optimally. The window may be short, but it is not nonexistent.
The capital threshold has moved far beyond traditional early-stage VC. 庄明浩’s market benchmark is that a vertical company reaching $100M ARR may receive a 30–50x valuation; Cursor is rumored to have raised $900M at a $9.9B valuation on roughly $500M ARR, despite its product version being only 1.0.
25. Multimodal models will consolidate, while vertical products can still charge for outcomes
潘乱 describes video models as “one general winning and countless casualties.” 6 or 7 companies, perhaps 7 or 8, may compete at the same time, half of them potentially large incumbents. 可灵, MiniMax’s 海螺, PixVerse, and others are still catching up, but the compute, data, and algorithm requirements for training are now far beyond the cost structure of traditional internet startups.
The voice market is more nuanced. 庄明浩 believes ElevenLabs has the strongest commercialization, with roughly $100M ARR; on leaderboards, MiniMax has performed best. But voice may not remain an independent battlefield for long, because video models are absorbing audio capabilities as part of the package.
His team uses a range of models, including MiniMax. It does not sell clients “whose tokens” they are buying, but delivers outcomes for clearly defined vertical scenarios and charges for the outcome. The technology sources can be mixed; customers are paying for completeness and reliable delivery.
美图 is the典型 example in 朋克周’s view. Its image and video models may not be the strongest, but it understands e-commerce, beauty, ID photos, and users’ willingness to pay, and has already generated solid revenue. HeyGen has also narrowed its focus to marketing videos, digital humans, lip-sync, and voice combinations. The story has shifted from “the strongest model” to a clear scenario, brand, and revenue stream.
26. The Agent era will rebuild an internet infrastructure designed for machines
庄明浩 observes that many early-stage US projects recently have stopped trying to build an end-user Agent and are instead filling in the infrastructure required for Agents to arrive. Existing websites and databases are built for people; AI identity, permissions, security, payments, and the ways APIs and SDKs are called may all need to be rebuilt.
This is also where the strengths China and the US have accumulated over the past decade-plus begin to converge. China spent years developing consumer mobile internet, products, and operations, while the US accumulated B2B SaaS, cloud, subscriptions, and enterprise delivery. AI products must understand enterprise willingness to pay while learning consumer-product experience; neither side can rely only on its existing strengths.
挨踢牛魔王 uses the steam engine to explain total demand. After the steam engine appeared, humanity did not reduce its coal consumption; more places began burning coal. As the cost of intelligence falls, it will similarly cover fragmented long-tail markets that were previously not worth writing dedicated software for, creating new demand for startups.
27. The most reliable conclusion is not to predict the winner, but to enter this round of organizational change
朋克周’s “honest non-answer” is: “Nobody knows what tomorrow will look like.” Google and Microsoft may not know either, while Altman and the founders of SSI may have marketing motives when they frequently discuss AGI. He has chosen to make media and observe from the shore, while respecting those who have already gone into the water to build applications.
潘乱’s practical reminder is that many domestic AI tools are still free. “Not using them would really mean leaving money on the table.” If the short-term result is not enough new value but mostly redistribution, people who are less skilled at using AI may lose efficiency and bargaining power to those who use it better.
庄明浩 places the long-term variables in OpenAI’s L4 “innovation” and L5 “organization.” In every industrial revolution, the appearance of new technology did not automatically produce a productivity surge; the real change came only after organizations and business models adapted. Google’s I/O is itself an example of organizational restructuring preceding product breakout.
He closes with Douglas Adams’s 3 laws of technology: technology that exists when you are born is ordinary life; technology that arrives between ages 15 and 35 is revolutionary; technology that appears after 35 is treated as contrary to nature. Although the guests are all over 35, they are still willing to embrace this change. “In a certain sense, we are lucky.”