Pioneers Insight Method Research Author
The Economics of AI Usage and What's Next For SaaS | Benedict Evans on a16z
Back to Episodes

The Economics of AI Usage and What's Next For SaaS | Benedict Evans on a16z

Summary

  • Agentic coding is AI’s first unmistakable product-market fit, while nearly every broader market-structure question remains unanswered. Evans says it went from “kind of useful to really changing everything,” with customers effectively pulling products out of vendors’ hands. But because this barely worked six months ago, predicting engineering-team design, junior hiring, or software careers three years out would be “insane.”

  • Foundation models appear structurally headed toward commodity economics unless their providers can prove durable differentiation or move up-stack. Evans sees no clear network effect, little differentiation beyond spending, and customers unlikely to care which model powers a SaaS product—just as they rarely ask which cloud hosts it. His deliberately hedged challenge is: the argument “deterministically looks like these things will be commodities,” so “explain to me why they won’t be.”

  • Today’s token economics are a transitory scarcity regime, not evidence of permanent pricing power. Users can receive “10 grand worth of tokens” for $20 or accidentally incur a $10,000 bill, echoing mobile data circa 2009–10. With perhaps $1 trillion–$2 trillion of capex arriving and models becoming “100x, 200x” more efficient annually, supply, usage, pricing, and ROI must find a different equilibrium.

  • The mobile-network precedent warns that vast usage and infrastructure spending need not translate into attractive returns. Mobile traffic rose roughly 1,500–2,000 times; networks collectively have about $1 trillion in revenue and spend around $200 billion annually on capex, yet their stocks have been flat for 20 years while “all the cool stuff got built by somebody else.” The central investor question is whether models become low-margin infrastructure or gain operating-system-like leverage—something Evans notes models currently lack.

  • AI is likely to create “way more software,” but that does not reveal which incumbent SaaS companies survive. Cheaper development, previously impossible functionality, and new combinations of probabilistic models with deterministic systems should expand supply and competition. Evans expects some percentage of SaaS companies to be wiped out, yet argues investors cannot identify them confidently enough to justify indiscriminately derating the entire sector by 50%.

  • The largest opportunities will come from making previously impossible products, not merely rebuilding old software with AI. Evans’s examples progress from finding a coat in an image to recommending alternatives and finally choosing one from a user’s Instagram that changes their look “but not too much”; enterprise systems might synthesize calls, emails, telemetry, and analytics to recommend pricing changes that improve churn. “The important stuff is not doing the old thing but more. It’s doing something new that you couldn’t have done with the old thing.”

  • Financial gravity will slow AI capex before technical possibility does, while much of the resulting value may be competed away as consumer surplus. Microsoft, Meta, and Google are each on course to spend more than 50% of revenue on capex, while the big four guide to roughly $700 billion collectively; Evans says the world simply cannot sustain $10 trillion a year of AI infrastructure. Even when AI turns a week-long DCF into a ten-second task, firms may perform 50 analyses without charging more—until today’s “magic” becomes something computers have seemingly always done.

Deep dive

1. Agentic coding crossed the product-market-fit line before the rest of AI

  • Evans’s biggest update since his previous presentation is that product strategies have diverged beyond “make a bigger model faster” and agentic coding has become the one use case with absolute product-market fit: “The customers are pulling it out of your hands.”

  • OpenAI appeared to try “everything all at once,” almost as though it had asked ChatGPT for 15 ways to build value above infrastructure and pursued all of them. Anthropic, with less capital, concentrated on coding and got it working—whether by deliberate strategy or discovery remains unclear.

  • Coding was directionally foreseeable because developers were the earliest users: the first thing people did with PCs was make computers, and the first thing they are doing with LLMs—“in a sense, LLMs are computers”—is make more compute. The timing and abrupt agentic breakthrough were not deterministically predictable.

  • Outside coding, adoption still ranges from Silicon Valley users running OpenClaw continuously on Mac Studio clusters to people saying, “Yeah, it’s kind of useful. I used it last week for something.” More prosaic enterprise demand is arriving one workflow at a time, such as a low-margin commodities company using LLMs to forecast when small producers will pay invoices.

  • ChatGPT numbers, model size, capex, and usage continue to grow, but Evans says open questions remain around whether there will be a model winner, whether models can capture value up the stack, how much they can do, and whether consumers will use them daily rather than weekly.

2. Six months of working technology cannot reveal a three-year labor market

  • Erik asked what coding agents mean for junior engineers, senior engineers, team organization, and careers. Evans’s blunt answer was, “I don’t think we’ve learned anything”—the technology did not work this way six months ago, and everyone is still scrambling to interpret it.

  • Automating work previously assigned to juniors makes old questions immediate: why were companies hiring junior people, and were they hiring them for the thing they did or for something else? Evans says those questions are real now rather than theoretical.

  • Evans rejected confident forecasts drawn from party anecdotes or current pricing distortions. It will take years for practices to settle, and anyone claiming to know what a software-engineering career looks like in three years would be “insane to think that you could know that yet.”

3. Mobile data shows how explosive demand can destroy infrastructure pricing power

  • Adoption comparisons must account for accumulated infrastructure. When Marc Andreessen was working on Netscape, there were only double-digit millions of PCs on the planet, so it could not produce 900 million weekly users; each successive platform inherits the devices, networks, and consumer habits built before it.

  • Early platform shifts also begin with unreliable products and a small group willing to put work into making them function. Turning that into a product where people can simply press a button takes time.

  • AI’s current pricing crunch resembles mobile data around 2009–10: some users received $5,000–$10,000 bills, while AT&T’s flat-rate iPhone plan encouraged 3G video until capacity buckled. Today, $20 can buy “10 grand worth of tokens,” while a few days of experimentation can generate a $10,000 invoice.

  • Mobile operators eventually aligned cost, price, and perceived value through caps, bundles, fair-use policies, and throttling. Traffic subsequently rose roughly 1,500–2,000 times; the industry now generates about $1 trillion of revenue and spends $200 billion annually on capex, yet its stocks have been flat for 20 years.

  • The uncomfortable analogy is that operators built transformative global infrastructure while value migrated upward. Evans’s core question for AI is whether models become commodity infrastructure sold near marginal cost or operating-system layers that “get to decide what gets built”; unlike Windows or iOS, models currently show no comparable network effect.

4. The commodity-model thesis is strong, but explicitly not certain

  • Evans sees no obvious way to make one model sustainably and fundamentally better than every other model. Products can emphasize different qualities, and users can prefer one, but the apparent competitive lever is mostly “your willingness to spend money,” not an Instagram-, YouTube-, or search-like network effect.

  • The chatbot is also a “weird, limited V1 UI.” Most professional tasks require the right data, configuration, controls, tooling, and interface; expertise in doing a job does not confer expertise in product design, just as excellent financial advisers are not necessarily the people who should design TurboTax.

  • Skills and templates may bridge some gaps, but users eventually outgrow them, just as departments outgrow Excel. Model labs cannot build every application any more than Microsoft or Apple could build every Windows or iPhone app, and enterprise buyers generally will not standardize on Claude or OpenAI inside third-party SaaS.

  • Evans imagines perhaps three to six frontier providers, plus edge and open-source models, spending an uncertain $200 billion–$2 trillion annually. Google’s advertising business also gives it different pricing incentives from OpenAI. Still, two surviving labs, model-level product absorption, or new up-stack leverage could invalidate his commodity case: historical comparisons “have no predictive value.”

5. The decisive AI questions are migrating out of Silicon Valley

  • Once a platform’s structure becomes obvious, Evans moves on: “The moment that you understand something and you know how it works and what’s going to happen is the moment you should move on to something else.” At this stage, “all bets are open,” much as Android’s apparent open-versus-closed victory over the iPhone once looked obvious.

  • AI’s impact on law, banking, consulting, advertising, and professional-services pyramids is increasingly an industry question. Evans compares it with Netflix: technology enabled the company, but the important decisions—shows, talent, awards, movies, and sports—became Los Angeles questions, not San Francisco questions.

  • Unlike earlier platform shifts, generative AI lacks knowable physical boundaries. Evans notes that people could look at their phones after the recording and find a new model priced at “2% of the price” because of an unexpected breakthrough, though he considers that unlikely.

  • Other open questions include whether older, open-source, or on-device models become good enough and how much compute can move onto devices. Because coding is currently the only clear product-market fit, another question is which field—law, banking, or something else—gets the next breakthrough.

  • Evans cited a run-rate move from roughly $9 billion to $47 billion, then stressed that the cited growth is all software: “That’s all software, isn’t it?”

6. Automation’s real payoff comes when it unlocks demand that never existed

  • Evans proposes several tests for each industry: does lower cost produce the same output for less, more output for the same, or still more output and spending—the Jevons-paradox question? Does it remove a barrier to entry, unlock a business model, or make a previously unimagined activity feasible?

  • Owning a printing press once protected newspapers; removing that cost changes competition. More radically, a steam engine enables a train that no quantity of horses could reproduce, while Spotify makes all recorded music available for $15 a month—an offer that was previously physically impossible.

  • Even a correct horizontal forecast can conceal industry-specific outcomes. The internet destroyed physical-distribution value, but that devastated newspapers while changing movie studios far less. “It depends” is not evasion here; it is the central mechanism.

  • Rebuilding the old product on the new platform is the paired fallacy: Google Docs reproduced Office on the web and won perhaps 20% share, but that was not the deepest opportunity. The exceptional startup “fills a hole in the universe”—often solving a problem the industry itself did not know existed, which is why a general-purpose model cannot simply “do the whole thing.”

7. Commerce illustrates the jump from correlation to higher-order intent

  • Advertising is roughly a $1 trillion market and retail $25 trillion, making better product understanding economically consequential. Google, Meta, and Amazon traditionally know SKUs, publisher metadata, and co-purchases—not why something is bought, hence recommendations that treat one toilet-seat purchase as the start of a collection.

  • Evans’s consumer progression starts with a coat photo: identify it and say where to buy it; then suggest ten similar coats at different prices with pros and cons; finally inspect Instagram and recommend a winter coat that changes the user’s look “but not too much.” Three years ago, the last step was science fiction; now it seems buildable.

  • The enterprise equivalent combines recorded Zoom calls, Salesforce email, product telemetry, and analytics to answer, “How should we change our prices to improve our churn?” That is a higher abstraction layer than sentiment-scoring angry callers, while Google and Facebook’s rising conversion and advertising metrics already reflect AI entering recommendation and prediction systems.

8. AI should multiply software while scrambling the SaaS hierarchy

  • Software will become cheaper and faster to build, previously impossible functions will become available, and competition will increase. The margin and pricing structure remains unresolved: outcome-based pricing sounds attractive, but tying each enterprise-software button press to P&L is often impractical.

  • Evans divides today’s enterprise fleet into big horizontal systems such as SAP and Workday, CRM, human capital management and payroll software; vertical applications; and an improvised middle of email, spreadsheets, and shared file systems. A large U.S. company may have 300–400 SaaS apps plus roughly 1,000 internally built or purchased applications running on-premises.

  • Scale determines the choice: PwC’s mass graduate recruitment warrants dedicated software, while a company hiring five graduates can use email and a shared Google Sheet. LLMs add more options—perform the task directly, extend Salesforce, augment a vertical app, or generate a bespoke tool that may become tomorrow’s mysterious 10-megabyte spreadsheet.

  • Models can sit at the bottom as a guarded Salesforce feature or at the top, synthesizing Salesforce, Workday, email, and analytics. The answer is probably both: probabilistic software handles interpretation while deterministic databases preserve reliable state. The conclusion is “more software—like, way more software,” even though an unknowable percentage of incumbents will be wiped out.

9. Tacit workflows and financial gravity constrain the transformation

  • Software companies and strategy consultants both inspect how a business operates and propose a better workflow; one encodes it in software, the other in processes, organization charts, training, and incentives. Much of the real operation is undocumented, including cases where employee bonuses reward ignoring the stated strategy.

  • This is why Bain, BCG, and McKinsey retain value: they can cross organizational boundaries, discover how work actually happens, and supply an external answer management can adopt or blame. Such tacit knowledge is difficult to package into a Claude skill that simply says, “Make a PowerPoint.”

  • Capital faces a harder boundary. Microsoft, Meta, and Google are each headed above 50% of revenue in capex, versus 15%–20% for telecom; the big four companies’ guidance totals $700 billion, compared with about $300 billion for telecom and $700 billion–$1 trillion for oil and gas. “We can’t spend $10 trillion a year” because that capital does not exist, while a frontier model may be relevant for only “3 to 6 months, 6 to 9 months, or whatever you want to say.”

  • ROI may surface as consumer surplus rather than higher profit: if a DCF falls from one week to ten seconds, analysts may run 50 without charging more; consultants may perform five times the analysis for the same fee and cost base. Evans closes with IBM’s early-1950s promise that an electronic calculator “gives you 150 extra engineers”: today’s magic will eventually become invisible—“Computers have always done that.”