The AI Industry Is No Longer Competing on Intelligence, But on Cost
Deep thoughts on AI and aspirations —— ByteDance Deep Thinking Circle
In 2025, Silicon Valley changed the axis by which it measures AI. In previous years, all discussions revolved around Scaling Law: whether models would continue to get smarter, when AGI would arrive. That year, the mainstream narrative shifted to token consumption: consumption continued to grow, rising over 20% month-over-month in July, with curves reminiscent of traffic and retention at the peak of mobile internet.
The strange thing is that in the same year, GPT-5 was released to mixed reviews—the model didn’t get noticeably smarter, yet usage continued to climb. Put these two facts together, and only one explanation holds: the industry changed arenas. Everyone stopped betting on the next intelligence leap and started managing the intelligence they already have.
Changing Axes Isn’t Decline, It’s Standard Operating Procedure for Industrialization
Why did expectations shift? Because demand was unlocked by existing capabilities. B2B wants cost reduction and efficiency gains, C2C wants to replace search and assist with work—these scenarios have long had sufficient intelligence requirements. What blocks users has never been models not being smart enough, but products not being smooth enough. As users switched from “waiting for the next smarter model” to “making use of what’s available,” metrics naturally shifted from capability to consumption.
GPT-5’s true identity becomes clearer against this backdrop: it wasn’t trying to prove itself smarter, but rather broke capabilities into separately priced modes like Instant and Thinking, and integrated dispersed models, information, and interfaces into a more usable product. This is textbook industrialization. Historically, electricity adoption didn’t depend on brighter light bulbs, but on continuously falling electricity prices; when performance converges, cost determines penetration speed. The “intelligence” metric itself has also failed. When models win gold and silver medals at the International Mathematical Olympiad, humans no longer have an appropriate yardstick to measure who’s smarter. Among the remaining competitive dimensions, cheapness is the only one that can still be measured accurately.
Companies at every position have therefore taken their places: model companies make each token more valuable, infrastructure companies make tokens faster and cheaper, and application companies figure out how to exchange consumed tokens for more data feedback. Agents particularly resemble Apps from the mobile internet era: before, each product needed an App; now, each scenario needs an Agent. And Agent token utilization efficiency is still very low—whoever gets efficiency up first captures the biggest piece in the middle.
Careful, Token Consumption May Be a New Vanity Metric
Here I need to pour cold water, and this is the most important independent judgment in this article: consumption is a cost-side number, not a value-side one. A hundred million tokens burned—what did it yield? Metrics no one questions will sooner or later become vanity metrics.
The mobile internet era went through an identical disease progression. First came traffic worship, with portals and Apps competing on install numbers; later, people learned to look at retention and ARPU, only then realizing half the traffic was false prosperity. The equivalent metric for the token era should be “revenue or retention per unit consumption”—in other words, the return rate per token. Judging industry health by how beautifully consumption curves rise is like judging mobile internet maturity in 2012 by download numbers.
Even more alarming is a term that became popular in Silicon Valley: Vibe Revenue. It means users know AI is useful but can’t articulate what exactly it can do, so much product revenue comes from inflated expectations and bandwagoning rather than real usage. This is a deeper sickness than ARR quality issues—ARR is at least real money; with Vibe Revenue, even users may not know what they’re buying.
The detection method is ready-made: don’t look at new additions, look at renewals. Whether the real task volume for the same cohort of users has increased over consecutive quarters, what percentage are paid renewals—these two curves don’t lie. Products whose revenue rises quickly but usage depth stays flat are likely in the Vibe Revenue zone.
Company Boundaries Are Disappearing, the Shell Layer Is Squeezed from Both Sides
The axis change brings a structural consequence: the division of labor between model companies and application companies has blurred. OpenAI recruited a batch of startup founders to do product; Cursor in turn started training its own models; Manus does both product and technology while proposing engineering methodologies like Context Engineering. Ambitious companies are all trying to integrate models, products, and infrastructure end-to-end, no longer accepting the old framework of “defining yourself in one sentence.” ByteDance is an App factory, Google is talent reserve—such formulations have failed in the AI era.
For the shell layer in the middle, this isn’t good news. Looking up, model companies casually build out scenarios with validated PMF, with each new model generation compressing shell value; looking down, inference optimization is largely open-sourced with questionable margins, leaving independent middle layers unstable. What can stand are the two ends: those with models, and those with scenarios and data. Companies in the middle with neither are seeing their survival space structurally narrow.
Primary market capital flows corroborate this. At the time in the US market, the valuation gap between first-tier and second-tier AI companies reached levels unseen in over a decade, with some star companies starting at valuations in the tens of billions. Heat multiplied tenfold, number of deals didn’t double, but money raised by invested companies multiplied tenfold. Money is concentrating at the top, and there’s no reason to see this trend reverse yet.
The Boundaries of This Assessment
I must be clear: the above views took shape in mid-2025, a mid-cycle judgment, not the end state. Two clearly defined verification points were identified at the time, worth checking item by item now. First, can Meta’s expensively assembled model team produce a significantly stronger model within 6 to 12 months; second, can the market produce application scenarios that stably consume tokens. If both give negative answers, the judgment that “AI valuations are broadly high” will start to materialize. The bubble here refers to overvaluation, not lack of technological promise.
For readers who want to track this continuously, watching three numbers is enough:
| Observation | What to Watch | What Signal Indicates Danger |
|---|---|---|
| Meta’s New Model | Whether it significantly exceeds same generation | Hiring without output, high-investment route disproven |
| Token Consumption Growth | Whether growth slows as base increases | Month-over-month growth declines for two consecutive quarters |
| Vibe Revenue Products | Usage depth during renewal period | High new revenue, renewal usage collapses |
One final sentence to close. The axis change itself isn’t bad news—when a technology moves from proving itself to deploying itself, metrics should naturally change. But the moment of changing metrics is most prone to muddying waters: in the intelligence era, storytelling is about capability; in the industrialization era, storytelling is about consumption, and consumption is precisely what’s easiest to inflate with subsidies and expectations. In judging an AI company, far more important than “how many tokens it burned” is what each token yielded. The former is an entry ticket, the latter is the answer to whether this business can actually work.