广密 on Coding as AGI's Second Act and Models as the Next-Gen OS
Summary
- 广密 sees the jump from Opus 4.5 to 4.6 as the quarter’s only generational leap worth highlighting: AI has moved from Chatbot’s first act into Agent’s second act, where it can complete high-value tasks. He sees the step as close to GPT-3 to GPT-4; model progress over the past quarter may even have exceeded all of 2025, while Miso and Spark would need to meet expectations to deliver a “true GPT-5 moment.” Another major leap could still arrive before June or July this year.
- Coding is no longer a vertical application but the second-largest accelerator of AGI after the GPU. At frontier labs, 70%-80% of a system’s code may still have been written by humans last year; this year, the share may already be below 1%. Claude and Codex are considered close to the level of a company CTO or Meta L8/L9 on many tasks, compressing the path from idea to working code from 2-3 weeks to 1-2 days or 2-3 days. “Language is the world, code is the solution” (语言即世界,代码即方案): if code can express most solutions in the digital world, much of the office white-collar workload could be automated.
- The core metric for AI businesses may shift from DAU to token usage among high-value users, with 1M-2M users at the top of the pyramid potentially worth more than 5M-6M subscribers. 广密 said Anthropic’s ARR has surpassed OpenAI’s and projected that the 2 companies could reach $80B-$100B by year-end and $150B-$200B or more next year; if that plays out, a teens-multiple PS valuation would not look unreasonable to him. The real constraint may not be demand but compute: “The biggest bottleneck to $100B of ARR this year may simply be compute.”
- Silicon Valley’s Big Three will continue to “each have their moment for 100 days,” with the 2026 hierarchy still far from settled. Anthropic’s strengths are top-down focus, data culture and a high-priced market at the top of the pyramid; OpenAI misread Coding because of ChatGPT’s overwhelming success but still has roughly a 50% chance of becoming the ultimate winner in 广密’s view; Google is 3-4 months late in the short term but has the strongest long-term foundation thanks to TPU, cash flow, Workspace and a systematic organization. “Today’s winning formula may be the poison of the next era.”
- The endgame for model companies is not a larger SaaS business but a next-generation operating system that supports an Agent ecosystem and automates global GDP. 广密 frames the roadmap as 3 acts—Chatbot, Coding Agent and Automated AI Researcher—and believes a company may announce AGI by year-end or early next year. His investment view is correspondingly extreme: if the 3-5 companies that keep delivering SOTA models ultimately automate 30%-50% of global GDP, each could reach a $10T valuation.
- White-collar deflation, an employment rupture and a painful window for America’s middle class may come before the technology boom. 广密 offered an aggressive but explicitly conditional estimate that “30% of jobs could disappear this year”; entry-level programmers, consultant demand, SaaS and IT outsourcing could all come under pressure, while even top AI researchers worry about losing their jobs in 1-2 years. Models make knowledge and intelligence cheap, while creativity, taste and IP may appreciate; his advice is blunt: “AI replaces people who do not embrace AI” (AI 取代的是不拥抱 AI 的人).
- Application opportunities remain, but the spillover window from stronger models will shorten; Harness and positive token ROI are the more durable frameworks. Products such as Cursor and Manus could be squeezed by upstream model companies building their own Agent products if they cannot access the strongest models; meanwhile, a good Harness can let ordinary or open-source models handle high-value tasks. For a One-person company, the hardest test is not a demo but the economic loop: “Spend $100 on tokens—can it make $110?” Outside models, 广密 is most bullish on robotics and AI for Science, which could see a step change over the next 6-18 months.
Deep dive
1. Opus 4.5 to 4.6 Pushed AI Past the Agent Threshold
广密 reduces the most fundamental change of the past quarter to a single event: Anthropic’s jump from Opus 4.5 to 4.6, a generational improvement close to GPT-3 to GPT-4. Models are no longer limited to chat, Q&A and web search; they are beginning to complete complex, high-value tasks, which is why he says task value and UC value are also rising.
His subjective impression is that progress over the past quarter “may have exceeded the progress of all of 2025.” If Anthropic’s Miso and OpenAI’s Spark meet expectations, they may deliver the “true GPT-5 moment”; in his view, the earlier GPT-5 repeatedly missed its mark, while Opus 4.5 and 4.6 looked more like half-generation advances.
Acceleration matters more than point-in-time capability: even if Miso and Spark are extremely strong, the next generation is already in training, and another GPT-3-to-GPT-4-style leap could still emerge before June or July. 广密 therefore believes “AI’s inflection point should already have arrived,” bringing the AGI timeline forward from 2-3 years to a possible announcement by year-end or early next year.
2. Frontier Researchers Are Becoming AI Managers, Not Coders
The most disruptive change at Silicon Valley’s leading labs is that frontier researchers and top programmers “basically no longer write code.” A system may have been 70%-80% human-written last year; the human share may be below 1% this year. The daily workflow increasingly resembles a professor or team lead supervising students: “AI writes, humans review,” even as humans begin to struggle to keep up with the review.
广密 says Claude and Codex are already close to the level of a company CTO, chief architect or Meta L8/L9 on many tasks, with a feature often becoming operational after 2-3 iterations. Top developers consume hundreds of dollars of tokens per day and thousands of dollars per week, reflecting not chat frequency but the fact that AI has entered real production workflows.
R&D cycles are compressing accordingly: the path from idea to working code that once took 2-3 weeks may now take only 1-2 days or 2-3 days. Data iterations for multimodal teams that previously took 1-2 months may shrink to several days or 1 week. As models become more efficient at processing data and building pipelines, they in turn accelerate other AI fields.
The stronger signal of a phase change is that “some breakthroughs in AI research are no longer coming from human engineers, but from Codex and Claude.” Anthropic’s release of 70-plus products and features in 50-plus working days is also seen by 广密 as a sample of organizational productivity that seemed almost impossible in the internet era.
3. “Language Is the World, Code Is the Solution” Explains Coding’s Generality
广密’s new strong view is: “Natural language describes the world; code describes the solution—language is the world, code is the solution” (自然语言是对世界的描述,code 是对 solution 的描述,就是语言即世界,代码即方案). Language and code are both highly compressed abstractions; if code can express most tasks in the digital world, most white-collar knowledge work performed on office computers shares an interface that can be automated.
张小珺 added that the generalizability of language and code has now been amply demonstrated, while mathematics may improve intelligence but can express a more limited set of things. Looking back at the previous season, she acknowledged that many people still treated Coding as a vertical use case and had not realized that it could turn a Chatbot into a real Agent that gets work done.
广密 divides the AGI process into 3 acts. Act 1 is the ChatGPT-style Chatbot; Act 2 is the Coding Agent, which directly completes tasks and accelerates AGI; Act 3 is the Automated AI Researcher, acting as everyone’s research assistant and pushing into foundational problems such as brain science, neuroscience and materials science. He even suggests that once Coding Agent matures, “90% of AGI may already have been achieved.”
Amazon’s book business is another layer of his analogy: books were simply the first SKU through which warehousing, logistics, users and supply chains were made to work, after which the model expanded horizontally. Coding may likewise be only a small part of the automation of global GDP, but it has the shortest and clearest feedback loop, making it both an accelerator and a testing ground.
4. No Leading Coding Model Means No Leading GPU
广密 believes a leading model company that neglects Coding will “most likely fall out of the first tier.” A proprietary internal model is also difficult to sustain because the task distribution within a single company is not broad enough to produce a leading system.
Reliance on Anthropic also creates supply risk for competitors: OpenAI has already been cut off, xAI has also been cut off, much of Google’s capability may be constrained, and Meta may not be safe in the future. His conclusion is direct: “Not having the leading coding model is like not having the leading GPU.” If one side uses A100 while another uses GB, the R&D gap will keep widening.
Coding is therefore neither a standalone industry nor merely a product; it is a means of production across the entire AI roadmap. It can multiply the productivity of the top 1%, or even the top 0.1%, by dozens of times and allow AI to accelerate AI itself, giving it a position just below the GPU.
5. The Revenue Flywheel Belongs to Top-Tier Token Users, Not Necessarily the Largest DAU
广密 says Anthropic has announced that its ARR exceeds OpenAI’s. More counterintuitively, the revenue contributed by Anthropic’s top 1M-2M users may exceed that of OpenAI’s 5M-6M subscribers. ChatGPT appears to have won the consumer market, but may not have won the higher-value Coding and Agent market, which “could be 10x to 100x larger.”
The yardstick for model platforms is changing. The internet era tracked DAU and advertising scale; the Agent era may focus more on usage, value and margin from super-developers and top-tier users. Coding reached in 2-3 years a run rate that took Google Cloud 17-18 years, an example 广密 uses to show that this revenue curve may be even steeper than ChatGPT’s initial breakout.
His projection is highly aggressive but remains conditional: if the momentum continues, OpenAI and Anthropic could reach $80B-$100B of ARR by year-end and $150B-$200B or more next year. At a teens-multiple PS, they would become the “Mag 7” companies of the new era. Anthropic’s biggest bottleneck may instead be compute; demand surged beyond its planning in the past quarter.
6. Coding’s Real Moat Is Organization and Data, Not Just Technical Know-How
张小珺 asked for Coding to be scored from 1 to 10: below 4, every company would do it well; above 8 or 9, Anthropic might continue to dominate alone. 广密 did not give a definitive score, but identified 2 decisive variables: organization and culture, and data.
Coding and Agent data are no longer ordinary text. They combine tasks, environments and evaluation systems, requiring extensive data generation, cleaning and validation. The problem is that the smartest people in a lab often want to pursue their own bets, chase zero-to-one breakthroughs and become the next Ilya, rather than spend years on the “dirty work” and “hard work.”
Management therefore has to choose the optimization target: consumer distribution in the style of ChatGPT, Gemini and 豆包, or Anthropic-style high-value tasks. Resources and attention are limited, and 2025 was a critical consumer window; Google and OpenAI were busy competing for traffic and, objectively, left Coding’s golden window to Anthropic.
7. Anthropic’s Lead Began as a Victory for Strategic Focus
Anthropic did not have Coding fully figured out on day 1. The consumer window had already closed, while the positive feedback from Sonnet 3.5 in summer 2024 gradually narrowed the company’s direction. It then went almost all-in on Coding, invested very little in multimodality, abandoned To C and did not follow the reasoning model trade that was being mythologized at the time.
广密 sees its advantage as top-down consistency. The founders understand the technology and personally push data and engineering; Jerry Kaplan is said to work directly with the data team. The idea that “the model is the application, and data is the model” appears embedded in the company’s DNA. Compared with OpenAI’s repeated shifts among pretrain, post-train, o1 and o3, Anthropic looks more like a team that has executed every industrial step well.
Team stability is also part of the outcome. Anthropic is not particularly interested in big names and prefers to hire underdogs; its culture interviews ask candidates how they would choose to live after AGI. Early offers were lower and the risks higher, yet those who joined were more mission-driven. Internal information is transparent, while leaks to the outside are tightly controlled, which helps explain why the company’s details remained opaque for so long.
Dario and Jerry Kaplan’s physics backgrounds shaped an approach of “observe the regularities, then scale.” They are not fixated on creating a new paradigm after Transformer, but continue to search for efficiency in data, architecture and engineering. 广密’s conclusion is that Anthropic has no secret formula; it simply executes strategy, organization, culture and detail all at once.
8. Claude Code’s Product Strategy Is to Capture Exponential Model Gains, Not Copy the IDE
After Cursor took off, the market broadly debated whether model companies should build another IDE. Anthropic instead had Boris, the founder of Claude Code, create a terminal-based product. 广密 sees this not as a feature trade-off but as an effort to “deliver the strongest way of coding,” allowing the product form to keep absorbing exponential improvements in the model.
Anthropic’s product team consists heavily of engineers and researchers who understand the model’s capability frontier more deeply, making them more efficient at converting capability into product experience. Claude Code’s Harness is also part of the Agent’s real-world performance.
The commercial positioning fits the technical strategy. Anthropic serves the top of the pyramid over the long term, allowing it to scale models larger and price them higher without rushing to cut prices, while preserving healthy margins. 张小珺 suggested that 70%-80% of revenue may be concentrated in Coding or Agentic products. 广密 acknowledged the concentration risk, but said the real defense still depends on Coding’s difficulty and the speed at which competitors catch up.
9. ChatGPT’s OpenAI Victory Temporarily Became a Liability in the Shift to Coding
OpenAI still has more than 900M weekly active users and 50M-60M paying users, although recent growth has been relatively flat. It has driven 2 paradigm innovations, and its overall strength, talent density and researcher status remain formidable. 张小珺 guessed that Coding only became OpenAI’s top priority in the past 2-3 months; 广密 believes OpenAI did make a strategic misjudgment on Coding.
广密’s central reflection is: “Today’s winning formula may be the poison of the next era.” Its massive free user base forced OpenAI to focus on inference costs and made it reluctant to build models that were too large. Anthropic, by contrast, used high prices, fewer users and high usage to support stronger models and secure the high-value market first.
GPT-5.4’s Coding capability has received strong community feedback and may not even be weaker than Claude’s, although its Agent capability still lags somewhat. 广密 sees this as more a question of time: OpenAI has no shortage of talent or resources, and a major strategic shift could be the starting point for correcting its mistake.
On market share, he offered only figures explicitly labeled as guesses. Sam said Codex had 3M active users; 广密 suspects that may mean weekly active users. Claude may have 15M-20M, putting the overall split at roughly 7:3. “I do not have the exact numbers; I am only guessing.”
10. OpenAI’s Freedom to Explore Creates Paradigms—and Resource Dispersion
广密 criticizes OpenAI’s biggest problem as a lack of focus. Sam’s VC background has made the organization resemble a broad capital-allocation exercise, with teams competing from the bottom up for resources. Multimodal initiatives such as Sora can be promoted internally into the main line, while GPUs are limited, multimodality consumes substantial compute and the direct benefit to OpenAI is unclear.
Culturally, OpenAI places particular value on zero-to-one breakthroughs but not equally on one-to-100 execution. Everyone wants to make a breakthrough, egos are strong, and data cleaning and product operations lack status. ChatGPT is enormous in scale, but 广密 says it “has no soul” and that it is difficult even to say who the real product owner is.
But the other side of the coin is that Coding now gives 1 or 2 people the ability to produce something momentous. OpenAI’s culture of free exploration could therefore generate the next paradigm-level breakthrough. 广密 gives OpenAI roughly a 50% chance of ultimately winning, believing that Anthropic’s advantage in incremental execution may not withstand another major technical leap.
He characterizes the 2 companies as pursuing different missions. OpenAI has always wanted to “be Einstein,” trying to leap over Coding and go directly to the AI scientist; Anthropic is first automating the entire white-collar world, with a more practical route and faster revenue acceleration. OpenAI has now turned its guns toward Coding, making it more likely that the 2 will continue advancing side by side.
11. Spark Could Be the Real GPT-5, but 2026 Is Still a Traveling Tournament
广密 expects Spark to be GPT-5 in the generational sense. In his view, GPT-5.1, 5.2 and 5.4 remained broadly at the GPT-4 level. Spark may not immediately surpass Anthropic in software engineering, where Anthropic has optimized deeply for Coding, but its other capabilities could move up across the board and form a new platform for continuous monthly iteration.
He favors a product direction that combines Chat, Code and Agent into an all-in-one platform. Over the long term, differentiation may depend increasingly on compute and iteration speed; he says OpenAI may have far more compute than “S O K.”
He rejects extrapolating the current leader to the full year: “The landscape has not stabilized in the past 3 years.” This year remains a continuing elimination tournament and traveling circuit. Scale effects, data flywheels and network effects once looked like walls against cold weapons, but models are beginning to iterate on themselves, and modern weapons may render old moats ineffective.
12. Gemini’s Short-Term Problem Is That Benchmark Wins Did Not Become Product Wins
广密 believes Gemini 3.0 was overestimated. Its benchmark scores were high, but real-world experience, sustained consumer growth and users’ willingness to pay did not keep pace, while its desktop product was incomplete. Its biggest achievement may have been a 2x move in Google’s stock and proving that the company was not an AI loser; Gemini 3 itself did not appear to represent a genuinely major breakthrough.
That success may even have reinforced the misread. Google continued prioritizing consumer distribution and multimodality, only making Coding the company-wide top priority 3-4 months later. Coding accelerates R&D speed, so “falling behind by 3 months may mean falling behind by 1 year”; catching up in the short term will therefore be difficult.
Over a longer horizon, Google may again be the most stable of the Big Three. TPU, cash flow, operating systems, Google Workspace and a complete organizational system are all in place; in the worst case, TPU could become another Nvidia. The company is already operating with its third generation of professional managers, more like a machine that can lose 1 or 2 people without disruption, making the probability of falling behind low.
Sustained delivery of strong models requires 3 things to hold simultaneously: investing tens of billions of dollars every year for 3-5 consecutive years; management with the judgment, courage and strategic conviction to make the bet; and a team capable of combining world-class talent with products. A one-off benchmark championship “is not very useful”; sustained delivery is the real test for every technology company.
13. Meta Is Now the No. 4 Seed, While xAI Is Held Back by Strategic Swings
广密 views Meta TBD as the challenger with the best odds, replacing xAI as Silicon Valley’s No. 4 seed. Its team has assembled know-how from multiple labs and produced a solid model in 9-10 months. Its roadmap is roughly 70%-80% modeled on Google and 20% on OpenAI, with Nano Banana in the corresponding direction and Mango for text-to-image and multimodal models.
Meta’s uncertainty is not limited to the model; it is also about how it delivers the product. 广密 leans toward a personal assistant, a personal friend or turning Open Cloud into a lower-barrier product. Chinese teams may be stronger than Meta at product innovation. High salaries buy time, but they may also attract people who stay only for the money and are unwilling to take risk, leaving team stability an open question.
Manus has not yet been visibly integrated into WhatsApp, Instagram or TBD after joining Meta, although revenue is growing rapidly. 广密 calls it the “ancestor” of independent teams building good Harnesses and says explicitly that it “was sold too cheaply.” But once Opus arrives in another 1-2 months, whether the standalone product could still command the same price will be unknown.
xAI’s problem is its repeated swing from large-scale pretraining to multimodality, Chatbot, AI search and then Coding. 广密 believes the bottleneck is data and data efficiency, not blindly increasing parameter counts. His analogy is “running a marathon at F1 speed, with the race taking place in a city”; without 200%-300% focus from the CEO and leadership, winning will be difficult.
14. Harness Engineering Is Building a Management Science for Agents
广密 recommends treating Agents as “first-class citizens.” Human knowledge workers have computers, work environments, permissions and even credit cards; Agents need a parallel infrastructure stack. In the future, tools may no longer be selected by humans but by Agents themselves, rewriting the definition of a software customer.
The other half beyond the model is the Harness. It resembles the organization, management and constraints a company provides its employees: normal people have a higher floor in a good organization, and ordinary models can handle high-value tasks inside a good Harness. When demand for Claude spills over and supply cannot keep up, non-SOTA and open-source models gain real room for use.
The commercial taxonomy is also shifting from To C and To B toward To Human and To Agent. For Agents, DAU may not be the key metric; token usage, the value of each token and margin deserve more attention. Token usage may also become a form of data, complementing pretraining data.
15. Models Are Becoming the Operating System for Automating Global GDP
广密’s endgame view is: “The model may be the next-generation operating system.” Everyday problems, work automation and scientific support will flow into a small number of leading models, whose infrastructure importance may ultimately exceed that of Google today.
The core function of an operating system is to support unlimited application expansion. In the model era, those applications are Agents. Models will create ecosystems resembling Windows, iOS, Android and WeChat, supporting different hardware including computers, phones and glasses, and ultimately becoming an “OS for global GDP.”
Chinese model companies are also converging on a direction. Over the past 3-6 months, Kimi, MiniMax and 智谱 have broadly shifted toward Anthropic-style Coding and high-value tasks because the consumer window is already difficult to compete in against 豆包. 豆包 leads on the consumer side, but it also “cannot afford to lose” in Coding and Agent; the competition will ultimately return to organizational capability and resources.
The bar for a new company entering this OS competition is approaching the challenge of rebuilding TSMC: invest $30B-$50B per year for 3-5 years, hire at least 100 world-class AI scientists, and secure GPUs, a strategic bet and go-to-market capability. New labs are not impossible, but until the technical path, data scaling and real scaling converge, where talent goes will say more than promotional claims.
16. White-Collar Deflation Will Create Pain Before Releasing Creativity
As knowledge and intelligence are compressed into models and converted into compute resources and tokens, the old social contract—read books, then exchange that knowledge for a job—is beginning to weaken. 广密 estimates that the social value of 70%-80% of people may change in subtle ways. ChatGPT and Claude have already reduced his need to buy consulting and other software; over the long term, many SaaS products may disappear, and Indian IT outsourcing may already have entered the model era.
Entry-level jobs will come under pressure first. He describes the US college-graduate employment rate as being at a historical low, with jobs requiring 2-3 or 3-4 years of experience already being automated. 张小珺 said Meta has laid off 16,000 people and worried that further cuts may follow; she also suggested that Microsoft may no longer need today’s 150,000 employees. The talent-development pipeline could therefore be cut in the middle.
广密’s aggressive scenario is that “30% of jobs could disappear this year.” Even the strongest AI researchers worry that their research workflow could be automated in 1-2 years, leaving only the next 1-2 years as a work and earning window. In a United States built on a middle class of more than 100M people, programmers, lawyers, doctors, brokers and bankers could all come under pressure, widening inequality and intensifying social tensions.
广密 still believes AI will ultimately bring prosperity, but human institutions have not kept pace with model speed: “AI’s intellectual progress over the past quarter may have been faster than humanity’s intellectual progress over the past 200 years.” After the painful window, creativity, taste and personal IP may become more valuable. 张小珺’s program outline was already largely generated by Claude Code from notes and an old framework, and the next season may be automated further.
17. The Investment Thesis Is Converging on SOTA Models, with Robotics and AI for Science as Satellites
广密’s investment conclusion is more extreme than last season: “Invest in companies that can keep building strong SOTA models.” If 3-5 model companies ultimately automate 30%-50% of global GDP, each could reach $10T, for a combined $30T-$50T. He spends 80%-90% of his own effort on models and suggests an investment vehicle directly expressed as a “model fund.”
Revenue data provides the most practical validation. Based on the figures cited on the show, though definitions may differ, Anthropic and OpenAI’s ARR are already above $30B and around $25B, respectively; Cursor is around $2.5B, Perplexity above $500M, Manus and Lovable above $400M, ElevenLabs and Suno above $300M, while Genspark is also growing rapidly. 广密 expects leading model companies to reach $80B-$100B by year-end.
Outside models, he prioritizes robotics and AI for Science, followed by infrastructure for the parallel Agent economy. Robotics could see a phase change in the next 6-18 months: an architectural breakthrough may emerge, technical paths may converge, and first-person and teleoperation data could drive subsequent data scaling and real scaling. Hardware is receiving renewed attention, with Silicon Valley teams beginning to recruit in Shenzhen—also a sign that Chinese teams may have an advantage in combining software and hardware.
One-person companies may still become the norm because models have sharply compressed the path from idea to working code to revenue. But the test for an application is a real economic loop: “You spent $100 on tokens—can you make $110?” His simplest work-level advice is: “AI replaces people who do not embrace AI; people who actively embrace AI may be the beneficiaries.”