Everyone Is Still Undersizing the AI Market | Eric Vishria
Everyone Is Still Undersizing the AI Market | Eric Vishria
Summary
- Vishria’s core macro call is that AI, like cloud, ends in an oligopoly plus hundred-billion-dollar specialists — not a single winner-take-all lab. In 2007 “you would’ve gone zero for thirty” polling smart investors on AWS’s margins; by 2014 the fear inverted to AWS eating everything, and both were wrong — Snowflake, Datadog, Cloudflare and a 40/30/20 cloud split emerged. His refrain for today’s “Anthropic’s going to do everything” anxiety: “What if it all works?” — while stressing most companies in each layer still fail, so differentiation matters more than ever.
- Inference is not a commodity: Fireworks runs the same open-source models on the same NVIDIA hardware as the hyperscalers at roughly 5X the speed plus a multiple-X throughput edge, paying the clouds’ margins and still making money. The investor takeaway — what looks from outside like a “pass-through resell” scale game turns out to conceal deep, specific expertise.
- The SaaS competitive frontier has moved, and incumbents’ choice is stark: “Get to AI or be worth three times revenue.” Database migration, once “the number one thing you would not do in software,” is now trivial because AI excels at well-specified interfaces and agents don’t tire — so cost, transportability and zero-to-infinity scaling become the arbiters, not old-fashioned migration stickiness. His most jarring board message: “every single day that you are hitting your plan, you are destroying equity value” — said to set founders free from pre-AI muscle memory.
- AI-native GTM breaks imported playbooks: proven sales leaders “completely flame out” because quota-capacity models assume pushed demand, while AI-first companies are “selling magic” with reps doing $10–30M; Patrick cites a $50M example. The best salesperson in these companies is the founder, “bridging the jagged edge to what the customer’s capability is.” Winning teams are those with three traits regardless of title: customer-problem understanding, taste, and curiosity about the “jagged edge” of model capability — building “sandcastles,” not castles.
- His biggest structural worry is energy, not distillation or Chinese open source: China is bringing on roughly ten times as much energy next year as the US. Since models “translate compute into intelligence” and demand for intelligence looks unlimited, less energy means fewer or more expensive tokens — his policy stance: solar, nuclear, gas, “do it all.”
- Cerebras taught him hardware’s brutal math — in software a sound block diagram gets you 80% of the way; in hardware “you’re like two percent of the way there” — yet he’s “astonishingly bullish” on AI silicon as the fifth great compute generation, and teases an unannounced investment in a sixth: a new CPU category for LLM-generated code. Each prior workload (CPU, graphics, networking, mobile) minted a hundred-billion-dollar company; AI’s new constraint is core-to-core communication.
- On venture craft: he invests in 1–2 companies a year (18 in 12 years) and screens with three questions — could I honestly convince someone I love this is their life’s work, would I pick up their 9pm Saturday call, and “if we’re right, will anyone care?” Luck framing via a banker’s golf analogy: “there is actually a way to increase your luck… get a lot of balls close to the pin. Eventually, one will drop.” The new growth fund reflects that high cash-on-cash multiples are no longer synonymous with early stage — “you could argue we’re a few years late.”
Deep dive
1. Fireworks’ lesson: running large models is genuinely hard, not a pass-through
- Vishria’s opening surprise from watching Fireworks: two-to-four-trillion-parameter models are “damn hard” to run, and running them efficiently is “super hard.” Everyone — AWS, Azure, GCP, neoclouds — serves the same stock open-source models on the same NVIDIA hardware, yet Fireworks shows “like five X” speed performance plus a multiple-X throughput difference visible only in the unit economics.
- The proof of moat: Fireworks-type companies are “paying the margins of the cloud providers and running on top and making money.” From outside, an investor sees a commodity “scale game” — “Oh, wait a minute. No, it turns out it isn’t. It isn’t at all.”
2. The AWS rhyme: zero-sum thinking failed twice, and it’s failing again
- The 2007 test: put thirty of the smartest investors in a room and ask whether AWS becomes a good business with durable margins — “I think you would’ve gone zero for thirty.” By 2014 the narrative had flipped to AWS eating all of enterprise (“you’re margining my opportunity”), and that was “massively wrong” too: Snowflake “out-Amazoning Amazon on Amazon,” Confluent, Elastic, Mongo, Databricks, Datadog at $100B — plus Azure and GCP going from irrelevant to an oligopoly he pegs at roughly a 40/30/20 split, with Cloudflare as a $100B player outside it.
- The takeaway he applies to “Anthropic’s going to do everything”: the market was so big one vendor couldn’t consume it. He’s explicit this isn’t spray-and-pray — “there was tons of roadkill,” relative winners matter — but AI is scaling even faster than cloud did, requires energy, power, shell, chips, memory and algorithms to build out, and “it feels to me like we’re going to end up with an oligopoly of winners” plus “hundred billion dollar crazy smaller winners.”
- He extends the non-zero-sum view across the stack: some CSPs, neoclouds, Fireworks-like inference providers, NVIDIA and some chip startups can all do well; edge inference on phones, near-edge inference on points of presence, and large models in data centers can coexist. That does not mean every company wins — most companies in each category will fail — but it does mean the value pool need not be divided among only one or two companies.
3. Enterprises want AI in a way they never wanted cloud — and “AI Sherpa” is the wedge
- The demand-side contrast: blue-chip enterprise was dismissive of cloud until roughly 2014–16; today enterprises with poorly absorbed AI are nonetheless running experiments, spending, and treating AI as both a bigger opportunity and a bigger threat than cloud. The Snapchat-at-40%-of-GCP era maps to Patrick’s comparison with Cursor’s outsized share of “all these things” — “same exact thing happened.”
- His advice to portfolio companies: the gap between Silicon Valley pace and enterprise adoption barriers is itself the opportunity — “let’s be their AI Sherpa,” crossing both worlds. Sierra’s Brett Taylor personifies it: start with automatable customer service, then extend to long-running agents with Horizon.
4. Sierra, sandcastles, and the inversion of product development
- Sierra’s setup matters: Peter and Brett have a 20-year relationship and are working on their third company together. They are technologists who have also lived in the enterprise world for a long time, giving them a basis for translating model capabilities into enterprise products.
- Sierra’s edge is being “very close to the metal of the models,” understanding “the jagged edge of AI capabilities, which is very different than the human smooth arc that we understand intuitively” — and building around it as new capabilities emerge “every four weeks.” Brett Taylor’s metaphor, which Vishria treats as the mindset test: “We used to be building castles. Now we’re building sandcastles and then they get washed away.” The artisan building a foundation to last a hundred years is “just not gonna make it.”
- Traditional product management — PM understands the customer, shields engineers from implementation — is now “a horrible way to do it.” The useful profile instead comprises “people who understand customer problems, people who have taste, and people who understand the jagged edge of AI capabilities and are curious about it. Those are the three things.” Patrick’s observation, which Vishria endorses: ironically “the returns to being technical are going up” even as models supposedly erase technical edge.
5. The SaaS frontier moved: hitting your plan now destroys value
- Vishria is “very dismissive” of the everyone-vibe-codes-their-own-software threat; the real issue is the competitive frontier shifted. His databases example: migration was software’s ultimate no-go, but Claude or Codex now builds against interfaces, AI is “very good at things that are very, very well specified,” and agents don’t tire of monotony — so migration becomes “kind of trivial,” and the winning criteria flip to cost, iteration speed, zero-to-infinity scaling and transportability.
- The message to SaaS CEOs: “Get to AI or be worth three times revenue” — brutal for companies that 4X’d since 2021 while multiples compressed from roughly 30X to roughly 6X (“multiple compression is a bitch”). More visceral still: “every single day that you are hitting your plan, you are destroying equity value” — a deliberate shock “to set them free” from a lifetime of plan-execute-compound training. Anne Lee Skates’ framing: CEOs run the core business 8-to-5 and do AI evenings, “and what they needed to be doing was the inverse.”
- The winners’ profile — Brendan from Recor, Lin at Fireworks, Brett, Max — are “so nimble about what the eval is,” constantly evolving the business. “That is very, very different than the way I was taught.”
6. Selling magic breaks the quota-capacity model
- Great sales leaders from the prior generation “completely flame out” at scaling AI companies. The mechanism: software sales was built on quota-capacity math — $1.2–1.5M early quotas, $2.5M enterprise quotas, discounted attainment — which implicitly assumes pushing demand. But “these companies are selling magic,” and if you’re selling magic first, “you’re gonna sell a lot more than two million” — reps doing $10–30M, with Patrick citing a $50M example.
- His hiring advice now: “check everything at the door… and just learn this from first principles.” And the best salesperson in these companies is the founder, whose job is “bridging the jagged edge to what the customer’s capability is.” The meta-error, same as cloud: “they just undersized the market… and this market is bigger.”
7. The binding constraint is energy — and China has ten times more coming
- His chain of logic: models “very effectively translate compute into intelligence”; intelligence demand looks unlimited; compute needs energy; and “China’s bringing on ten times as much energy next year as we are in the US.” Therefore less US energy means “a lot less intelligence or a lot less tokens or a lot more expensive tokens… and that seems very bad.” He rates this above distillation and Chinese open-source models as the real worry.
- The prescription is deliberately agnostic — he mentions gas turbines, natural gas, solar and rare earths among the relevant areas: “We should do it all, and it’ll work itself out.” He won’t pick the relative winner: “I’m not smart enough to predict which one’s which, but it’ll all work.”
8. Cerebras: naivete, roof lines, and a sixth compute generation
- The 2016 pitch — five founders and a deck, pre-transformer, NVIDIA at $40B not $4T — opened with “GPUs actually suck for deep learning. They just happen to be a hundred times better than CPUs.” Only three levers were identified to speed up deep learning in hardware: more cores, more core-to-core communication, and memory closer to compute — and Cerebras took all three “to their logical maximum”: a wafer-scale chip, roughly 450,000 cores, and roughly 20GB of on-chip SRAM.
- The hard lesson: in software a sound logical block diagram gets you “eighty percent of the way there… In hardware, you’re like two percent.” Sims set your roof line — “best it’s ever gonna be” — then compilers and kernels take away from it; engineers may start around 10% and grind upward for months and years. Physics and a supply chain involving TSMC and roughly 30 other vendors also matter. The visceral moment: a 2019 board meeting with roughly $500M raised and the part “fucking melting” — “I was just like, we’re gonna lose all this money.” Asked if the experience makes him want more hardware deals: “Fuck no” — though he mentions a robotics company and an unannounced CPU investment anyway.
- The framework that justified it: each new workload — multipurpose compute, graphics, networking, mobile — produced a new $100B company (Intel, NVIDIA, Broadcom/Avago, Qualcomm/Arm). AI’s new constraint was core-to-core communication. And he sees a sixth generation coming: LLMs on accelerators generate code that runs on classic CPUs carrying old baggage — “there’s actually room for a new CPU approach.”
- On when naivete is a bridge too far: the stereotypical objections are right “nineteen times out of twenty. Maybe ninety-nine times out of a hundred.” Bruce’s question — “What could go right?” — plus a saying relayed to one of his partners: “If it doesn’t work, it’ll be for all of the reasons that your partner said. If it does work, it will be because those reasons didn’t matter.”
9. Robotics: bootstrap the data flywheel; the task is almost irrelevant
- Classic robotics has largely solved repetitive tasks in controlled environments; the harder opportunity is unstructured tasks in real-world environments, where the robot needs AI. The core problem is that “there is no internet scale data to bootstrap the whole thing” — LLMs had the internet, robots don’t. The insight he likes: prioritize high-value data rather than slop, build a pre-trained base, then add small auxiliary post-training examples for new tasks — the same pre-training-plus-RL magic as LLMs. Sunday Robotics runs this pipeline with gloves designed to match the robot’s hands for clean data transferability; his visceral moment was descending from a “totally janky cardboard glove” Stanford-basement demo to a dozen robots folding arbitrary laundry with rigorous measurement of fold quality — “Oh my God, this is happening.”
- To Patrick’s worry about solutions in search of problems, Vishria doesn’t share it: laundry is a good task precisely because it’s arbitrary, dexterous, and time-insensitive, but “I don’t think the task is actually that important… If you get that flywheel going, then the task capability will just keep multiplying.” The AV analogy: Waymo and Tesla both vertically integrated robot, model and data collection — a simplification that ships a complete product or solution and, he says, should ultimately generalize in some way.
10. The craft: partner-first investing, IPO windows, and the Hinton error
- Why founders and other investors name him a strong board partner: “I’m an investor second, and I try to be a partner first.” He’ll pass on “investment-grade” deals that would make money absent chemistry; his screens: could he honestly convince someone he cares about that this is their life’s work; the green-button test (“if this person calls me at 9:00 p.m. on a Saturday night, will I pick up the phone?”); and “if we’re right, will anyone care?” The Benchling grind — “they got seven years of churn in like twelve months” post-biotech crash — is his exhibit for staying through it.
- The partnership is based on mutual selection and repeated learning: the founder must want to work with him, and he wants to learn from and help the founder. Much of his job is asking questions that help founders solidify their own conviction; a few 1% or 2% better decisions compounded over a decade can produce real results.
- On luck, the banker’s golf framing he uses with his kids: keep getting balls close to the pin; “Getting the hole in one, that’s luck.” Work with special people on unbounded opportunities, focus on what the company can control, and accept that timing, macro conditions and supply chains can still determine outcomes.
- The growth-fund logic: LPs invested in venture for the possibility of exceptional multiples, not merely to beat an index by a few percentage points. For most of venture history, early stage and high cash-on-cash multiples were “synonymous… those two circles in the Venn diagram almost perfectly overlap.” Bigger outcomes broke that overlap — not a “gazillion” hundred-Xs outside early stage, but enough. He concedes “you could argue we’re a few years late,” and says the real blocker had been having a team built for it; they had passed on right intuitions “because it was outside the box. That’s obviously dumb.”
- The partnership also debates where AI value accrues across infrastructure, applications, foundational models and semiconductors. Vishria expects business-model innovation such as selling by outcome, analogous to SaaS subscriptions, to matter alongside the technology.
- On going public: currency, capital access and trust — he thinks the labs will ultimately go public and that it is “beneficial to the world and to America” for transparency; his athlete analogy: what does a collegiate athlete want? “Go pro.” The warning: windows close — roughly 500 private SaaS companies between $100M and $500M whose employees and investors are “stuck” because AI natives “sucked all the oxygen out of the room and… the window was missed.”
- His closing parable against deterministic doom: Geoffrey Hinton, “three orders of magnitude smarter than I am,” said, perhaps in 2016, to stop training radiologists — “could not have been more wrong” despite correct premises, because the aggregated training data didn’t exist, reimbursement and malpractice liability complicate substitution, and Jevons-paradox imaging growth means “we need more radiologists, not less” during a co-pilot phase that may last a long time. New Lantern plays here. His verdict on mass-unemployment predictions: “it’s almost the exact same setup.”