Nikesh Arora on the Frontier Model Problem: Breadth vs Depth | The Future of Token Costs
Nikesh Arora on the Frontier Model Problem: Breadth vs Depth | The Future of Token Costs
Summary
- The headline call: token prices fall, with reductions over 3-5 years, toward 1/10 of today’s level. Arora’s mechanism: compute is scarce and costs “two to three X or four X more than it used to cost 2 years ago,” half of it feeds free consumer usage that is “fundamentally loss-making,” so the pressure lands on the paying half — enterprise coding. Once frontier labs have enough post-training data, he expects them to actively constrain consumer use, and cheaper tokens then unlock consumption — which is why he won’t call whether token spend per developer settles at “3.8 or 20” percent of salaries.
- The frontier model problem is breadth vs depth. Consumers are “highly tolerant” of false positives — Gemini wrote him a passable investment memo in 4 minutes — but enterprise agents making independent decisions demand zero tolerance. Waymo is his benchmark: “the biggest agentic product out there,” built on tens of billions of dollars of edge-case training on data that isn’t on the internet. “You can’t stick the next model of Anthropic into your Mercedes and say, ‘Okay, drive me home.’” Coding is the one universal enterprise use case; frontier revenue beyond it needs depth.
- Memory becomes the moat — and the fight is over where it lives. Frontier models “will spend a lot more time in the next year or two building memory around consumption,” because user context creates stickiness. The risk for enterprises: memory embedded in the model makes you “model captive, not model agnostic,” and the orchestration layers that could keep you agnostic “are not as well funded as these models.”
- Enterprise AI is not ready, and most buyers are doing it wrong. More than half of enterprises are bolting AI onto existing workflows for 20% gains instead of rethinking them; SaaS “has no opinion. AI applications will have opinions.” His rule of thumb: half the people in G&A functions (marketing, finance, HR) within 3 years — but more technical and more sales headcount, not fewer people overall.
- Mythos is “an accelerant to cybersecurity,” not a threat to it. Pointed at Palo Alto’s own code, it found in 6 weeks what would have taken 5-6 years — but a defensive auto-patch “is going to patch 30% things which are not wrong,” so attackers get weaponized faster than defenders get automated. That asymmetry lights a fire under security budgets, and he says there is no “cloud” or OpenAI endpoint agent to displace his 150 million sensors at the gate.
- The SaaS repricing is rational confusion, not oversold panic. Systems of record face workflow reimagination, the analytics layer is being absorbed by data lakes plus LLMs (Snowflake, Glean, Databricks), and nobody knows how many seats survive — “I don’t know what the right valuation is.” On Salesforce’s best days: “that depends on how they execute from here on.”
- Consumer AI may not be ad-funded at today’s compute cost. From his Google CMO years: online already took 60-70% of a ~$500-600B ad pie growing only 3-5% a year, so ad dollars for AI must cannibalize existing budgets. The opportunity is transaction revenue — average conversion runs 1.5-2%, meaning “85 to 90% of marketing is wasted,” and consumer goods carry 92% distribution-and-marketing cost that AI targeting can compress.
- Quick-fire alpha: FOMO has compressed diligence windows — “You had 20 years to invest in SpaceX. You had three to invest in Anthropic” — and the discipline against it is his board member’s sunk-cost walk: ignore the 3 months of effort, ask “if this walked in the door right now… would I take it or not?”
Deep dive
1. Breadth vs depth: consumers forgive false positives, agents can’t have any
- Arora’s framing of his own tweet: frontier labs keep leapfrogging each other, but “even the best model has a high false positive rate” — and the consumer doesn’t care. There’s always a human in the middle exercising judgment; his sister happily quizzes ChatGPT, and Gemini produced him an investment memorandum in 4 minutes that “would have taken me days” with bankers and analysts. That breadth makes a frontier model the go-to consumer brand, with all the distribution economics of YouTube or Google search.
- The enterprise is the opposite regime: “if you imagine a future where an agent’s going to make independent decisions and act on it, you have zero tolerance for false positives.” His benchmark is Waymo — “the biggest agentic product that is out there” because it fully replaced a human — built on “tens of billions of dollars” of edge-case training and proprietary data. “You can’t stick the next model of Anthropic into your Mercedes and say, ‘Okay, drive me home.’”
- The structural tension: frontier models chase consumer attention because it drives post-training data and brand, while “the real enterprise revenue is going to come from use cases that require a lot more context.” Coding is the one standout — a universal activity where everybody’s data helps train the model. “Hopefully there’s a few more out there.”
2. Enterprise AI isn’t ready — and software is about to grow opinions
- His diagnosis: “more than half the enterprises are still not getting it right” — grafting AI onto existing processes for marginal gains (“look at that, it’s happening 20% faster”) instead of rethinking workflows. “The winners in the long term will be people who actually rethink their companies with AI, not people who adapt their current workflows marginally.”
- The shift he describes: SaaS containerized common processes but “the software is not intelligent.” The next generation judges — screen every CV, pick the 20 to interview, tell Harry which 10 questions to ask because three colleagues only asked five. “That requires us to give up human control and let AI do 80% of the thinking for us. That’s not how we’re doing it right now.”
- His formulation worth keeping: “SaaS applications will give way to AI applications. The difference being SaaS applications have no opinion. AI applications will have opinions” — an AI marketing assistant that says “I looked at your copy, it sucks… it’s not consistent with tone of voice.”
- The headcount math: marketing is the best-trained domain because marketing is by definition public data, so why 600 people in marketing? Rule of thumb: “in the next 3 years we’ll probably have half the people in G&A type activities” — but he pushes back on the jobs-apocalypse read: “I think we need more, not less” technical resources (prompting, harnesses, proprietary data) and more salespeople — after 20 years, half the European customers he met last week still don’t know Palo Alto’s full catalog.
3. The Darwinian workforce: hackathon hiring and the token whack-a-mole trap
- The constraint: “90% of the enterprise employees are not AI-savvy… There’s no course you can take at any school anywhere. They have to learn on their own.” Two paths out: the Brian Armstrong / Jack Dorsey model — “decimate my organization,” rebuild at 30-40% headcount because “there’s no redemption” — or the gradual one. Palo Alto chose gradual: hiring only through hackathons, replacing ~2%-a-month natural attrition with AI-savvy people. “Give me 12 months, I’ll have transformed 20, 25% of my team.”
- On token budgets, his warning is the sleeper insight of the episode: “your smartest employee who knows how to use the AI really well could be using 20 times the tokens that an average employee uses. If you get into this whack-a-mole moment saying, ‘I’m going to stop people spending too many tokens,’ you actually will hurt the best AI-savvy people more than the average employee.”
- Harry’s clarifying push — is it a free-for-all? — draws the actual policy: “we have a use it judiciously model… if we find somebody who’s using it well, we won’t constrain them.” Build only what’s proprietary; where a generic AI application will exist in 12-24 months, “let’s just wait.”
4. Token economics: prices should reach 1/10, because the consumer half of compute makes no return
- The mechanism, spelled out: “there’s not enough compute for what the world is demanding. Unequivocally.” Compute costs 2-4x what it did two years ago, and “more than half of the compute is going to feed the consumer, which is a fundamentally loss-making entity right now” — free, billions of users, no return. “Guess where the pressure goes. The pressure goes on the other half of compute” — enterprise coding — which must pay until consumer transaction or ad models mature. Unlike the Google/YouTube era, “the compute requirements and the cost is now 10X” of what they were then.
- The call: “I think the long-term token pricing should be 1/10 of what it is today,” with reductions over the next 3-5 years — which is why he won’t adjudicate Harry’s framing (a recent guest’s $300M/year Anthropic spend equalling 3.8% of developer salaries, vs. “Brandon at Macquarie” saying token spend matches salaries): “pricing will move very drastically… not sure we can tell the answer right now.” He also expects labs to constrain consumer usage once they have “enough post-training data, more than they need.”
- Why prices haven’t fallen yet: “your frontier model companies are value maxing, not token maxing.” At a trillion-dollar valuation needing another hundred billion of compute yearly, “the only lever they have is to take the fastest growing thing… and charge us more for it.” And the capability excuse doesn’t hold: “I don’t need Fable 5 or Mythos 5 to do 90% of what people do with the AI today” — Harry concurs that two-year-old models handled most queries; Arora’s caveat is they were compute-inefficient, and the R&D bill is being recovered through tokens.
- On the infra bubble question, an honest open loop: “I have a question which I don’t know the answer to… at what point in time does physics kick in?” Data centers, energy, copper — if capacity buildout outruns physics-limited deployment, “there may be a digestion period,” which changes timing, not the need: “we will still want as much compute as we can deliver as fast as we can deliver.”
5. The advertising pie may not fund consumer AI — transactions could
- Credentialed math from his Google days: in 2004 Google was 2% of global advertising, a pie of $500-600 billion; online is now ~70% of total ad revenue and the pie grows maybe 3-5% a year. “That money that you’re hoping to fund consumer AI from advertising will have to come from current advertising revenues” — so an OpenAI ad engine doesn’t “change the equation drastically to make the consumer profitable.”
- A possible bigger pool is transaction revenue. Harry guesses online conversion at 1-1.5%; Arora: best of breed is actually 7-10%, average 1.5-2% — “which means 85 to 90% of marketing is wasted.” Consumer goods cost 5-8% of list price; “the 92% is distribution and marketing.” AI with memory and context that targets Harry at the moment of purchase intent pulls traditional marketing dollars online as transaction take, not ads.
6. Memory is the moat — and it decides whether you’re model-agnostic or model-captive
- His crystal-ball call: “the frontier AI models… will spend a lot more time in the next year or two building memory around consumption” — more than context windows: storing 30/60/90 days of your interactions so every answer compounds. “As you start building context in a user base, you create stickiness, and that becomes your moat.” Applications with memory exist in consumer; “that has not yet come to the enterprise space, funnily enough” — when it does, enterprise demand for compute and memory rises again.
- The architectural question investors should sit with: “do I store my memory and context in the orchestration layer or in the frontier model?” The models know this is the moat and are moving aggressively; orchestration layers “are not as well funded as these models.” The risk: “you end up in an architecture where… you cannot be model agnostic, you actually become model captive to get maximum efficacy” — switching means redesigning your entire application.
- On Chinese open source, he flips Harry’s question with a thought experiment — take the word China out: open source models aren’t dangerous, so the real worry is nation-state backdoors: “does the model wake up one morning and it’s got a sleeper agent in it and starts sending all the data somewhere else? Those can be secured.” He sees the world bifurcating into task-specific depth models (ElevenLabs in voice; physical AI for planes ≠ cars ≠ robots, “a depth use case only”) under smarter orchestration layers — and open source “allows you to play the cost curve. I don’t need the smartest model to do the smartest thing.”
7. Mythos is an accelerant to cybersecurity, not a replacement for it
- The offense/defense asymmetry: trained on great code, the model “also knows how to find bad code” — pointed outside-in at an enterprise it finds open web sockets and misconfigurations that attackers can daisy-chain. But defense can’t be automated the same way: “it’s going to patch 30% things which are not wrong. Who knows what that’s going to do to blow up your infrastructure.”
- The internal proof point: Palo Alto ran Mythos against its own code and “found in 6 weeks what would have taken us 5 to 6 years” — but patching still required human evals, sandbox and production testing before shipping. Net effect: “it’s lit a fire under the security practitioners around the world… it creates urgency on the part of customers to improve their cybersecurity posture, which is generally a good thing for cybersecurity companies.”
- Why he isn’t disintermediated: perimeter defense needs someone at the gate — “we have 150 million sensors in the world… there’s no ‘cloud’ endpoint agent, no OpenAI endpoint agent that I can replace Palo Alto with.” The post-breach problem — finding the intruder fast — is the true AI task, needing exactly the enterprise context he’s spent 5 years building. On government intervention: guardrails “have not been built robustly enough” and jailbreaking proves it; is fixing them possible? “I’m hoping it is.”
8. Transform like Tesla, not like a carmaker AI-washing its lineup
- Asked what he’d do differently starting today, he maps three self-driving strategies onto enterprise product: Waymo (bound and train until full autonomy), Tesla (automate segments, human takes the edge cases, FSD improves), and traditional manufacturers “trying to stick a little bit of AI and sort of AI-washing their cars.” His verdict: an incumbent with customers to satisfy “has to have the Tesla approach” — you can’t ship a product that’s “right 80% of the time” and ask for 3 years of patience. His stated fear: “am I pivoting fast enough?”
- Why he can’t do the Armstrong/Dorsey implosion: “I don’t think the underlying application software is there.” He refuses to build generic AI marketing/HR/ERP stacks — “I’m hoping somebody goes and does that… perhaps the next iteration of Salesforce, SAP, or Workday” — and builds only where Palo Alto has unique context and memory.
- The operating cadence: a twice-weekly meeting called AI EIEIO (“like Old MacDonald had a farm — everybody in my company wants to do AI”) where the top 15-20 technical leaders demo what they’ve done. The design is deliberately Darwinian: “whatever motivates you — the fear of Nikesh asking you in 3 days what you did, or your team pushing you — you will show up with something.” Transformation runs top-down, not bottom-up; bottom-up token experimentation just surfaces the best talent.
- The trap he’s avoiding, seen before: 2004’s “chief internet officers” — 24-year-old sherpas CEOs hired so they could wash their hands of the internet. “The risk of that happening is true with AI as well… meet my chief AI officer, who was probably a researcher at some amazing university and has low execution skills.” Product roadmaps prove the drift: 6-12 months out, “there’s nothing called agents in there.”
9. FDEs are an honest confession that the product isn’t finished
- Harry stages the fight — Factory’s [likely Matan] (“if you need FTEs, you have a shit product”) vs. the Palantir school — and Arora refuses the binary: enterprise AI has been chased “for the last, what, 12 months? At best.” His translation: “FTE is a short form for saying ‘my product’s not fully there… I’m going to send some people who are going to sit in your office and build my product while I adapt to your needs.’” A real FDE brings code back into the product; the rest are technical sales consultants. Startups are hungry for revenue and being told to “sell it before the product is fully ready” — so FDEs are needed, short term.
- The churn evidence from coding: Wind Surf and Devin were the early names — “they don’t exist in their then form” — now it’s Codex, Claude, Antigravity, with Factory and Cognition on SDLC. “Who knows in 2 or 3 years who’s going to be the leader.” Will he bet? “No, you do that. I just need a good one that my teams can use.”
- His M&A logic in the same spirit: he bought an agentic-AI gateway company 6 months ago, cheap, on the thesis that governing enterprise agents requires aggregating agent traffic through a gateway you can watch and stop. Now people are waking up to routers for optimization and token reasons — he might have paid double later, and it wouldn’t matter: “things I buy either help me 10x or 100x, or they fail spectacularly. The one-to-two-x doesn’t make a difference.” The paranoia underneath: “you miss one trick, you can survive. You miss two tricks, you’re partly impaled. You miss three tricks, you could be obsolete.”
10. The SaaS repricing, the PANW bull/bear, and the discipline against FOMO
- On whether SaaS is oversold, he decomposes the market’s message into three stacked confusions: systems of record face workflow reimagination (opinion-less software giving way to opinionated software); the analytics layer sitting on top is “getting reshaped already” as Snowflake, Glean and Databricks let LLMs run against data lakes; and nobody knows how many seats survive. Honest bottom line: “I don’t know what the right valuation is.” On Neill Mehta’s test — are Salesforce’s best days ahead or behind? — “That depends on how they execute from here on.”
- His own bear case, unprompted and clean: “if I can’t make that transition happen with my team in the next 3 years, yes, there’s a bear case because somebody else will build a better mousetrap.” The bull case rides platformization — customers refusing to manage 40-60 security vendors — plus runway: Palo Alto went from under 2% to “closing in on 8 or 9%” of cybersecurity revenue, leaving room to 20-40%. Venture-scale returns survive because attackers innovate — but “a billion dollars doesn’t do it anymore. It needs to be 10… 20, 30.” And the macro kicker: marketing tech, HR tech, and token spend replacing repetitive human work “all becomes tech spend” — tech’s share of the S&P keeps rising.
- The quick-fire take on Silicon Valley’s consensus error: euphoria plus FOMO, on compressed clocks — “You had 20 years to invest in SpaceX. You had three to invest in Anthropic” — breeding the belief that “every company that’s going to show up now is going to be the next Anthropic, so we better get into it.”
- The antidote, from his own board on a near-$1B acquisition: “sometimes you confuse effort with wanting the outcome… You haven’t spent a dollar yet. You just put in 3 months of effort.” His practice now — and his advice to VCs who beat eight rivals to a term sheet: go for a long walk and ask, “if this walked in the door right now, and there was zero effort involved, would I take it or not?”