Pioneers Insight Method Research Author
The Early Days of Anthropic & How 21 of 22 VCs Rejected It | The Four Bottlenecks in AI | Anj Midha
Back to Episodes

The Early Days of Anthropic & How 21 of 22 VCs Rejected It | The Four Bottlenecks in AI | Anj Midha

Summary

  • Scaling laws are alive — saturation is a property of the domain, not the paradigm. Pushing back on Demis’s diminishing-returns framing, Anj Midha is categorical: coding evals can be saturated, but at Periodic Labs (his latest incubation — LLMs predict superconductors, robots synthesize them, X-ray diffraction verifies, data pipes back into training) “throwing more compute at the problem is probably having super exponential gains right now per iteration… The bitter lesson is holding, well and alive.”
  • The Anthropic seed was rejected by 21 of 22 investors — one asked “What’s GPT-3?” of the team that invented it. The raise re-anchored from $500M to a $100M seed while OpenAI had raised $1B; the “compute multiplier” pitch (a unit of intelligence for 6x less per VC dollar) was legible only to EA-adjacent ML people like SBF and to Amazon, which did the original $4B compute-for-equity deal. Zero VC firms were in the seed; Midha put in his life savings and has since deployed “many hundreds of millions” into “the world’s fastest-growing business of all time.”
  • “We are not in an AI bubble… We are definitely in a GPU wastage bubble” — billions of dollars of stranded compute sits idle because flops aren’t fungible across H100s, GB200s and GB300s. We’re in compute’s pre-standardization era, “1885 industrial revolution England,” and if OpenAI, Anthropic or Gemini miss revenue targets in coming years, his stated reason will be compute access, not demand.
  • China’s full-stack race is not a chip race: co-design Huawei silicon with infrastructure and training runs, adversarially distill Western frontier models at scale, release open models to bootstrap feedback — then stop open-sourcing once caught up. His answer is a Western “iron dome” for inference: “If we don’t secure frontier model inference… behind a coordinated iron dome, I don’t think we have a sustainable shot at staying at the frontier over the next decade.”
  • The Cloud Act is the first crack in hyperscaler dominance in 15 years — mission-critical European workloads (ASML, CMA CGM) legally can’t run on American-managed clouds, which is his entire Mistral thesis: “independence at scale of every part of the AI infrastructure stack.” Europe’s sovereignty bar is Google-level infrastructure — roughly 12–15 gigawatts — over the next 4 years.
  • “Perfect competition is for losers” — and monopolies are mafias. His update to Thiel: the optimal market structure is 3–4 teams per frontier; the ~50 VC-subsidized inference companies are “lighting hundreds of millions of dollars on fire,” and the 4–5 winners will be separated by one variable: “Supply. Access to supply… If you’re making a steam engine, you need coal.”
  • Venture is going “back to the future” — Arthur Rock at Intel, Genentech in Kleiner’s basement, Markkula at Apple — with value accruing to investor-cofounders, not check-writers. Midha’s broader investor method is to treat the future as undetermined, form bottleneck hypotheses, run parallel experiments, and stay willing to be wrong. Amp is built as the “independent system operator” of a compute grid: 1.3GW secured as proof of concept, ~$40 of cloud spend over 4 years, financed ~20% equity / remaining debt, with compute given away at cost under public-benefit governance. His LP advice: “I would be investing in the bottlenecks.”

Deep dive

1. Scaling isn’t saturating — saturation lives in the domain, not the paradigm

  • Harry opens with Demis’s suggestion that returns to compute are diminishing. Midha’s rebuttal is unhedged: “Oh no, absolutely not… that’s not true at all.” In well-explored domains like coding, yes, incremental eval gains cost more compute — but that’s the saturated exception, not the rule.
  • The example that carries it: Periodic Labs, his latest incubation, a 30,000 sq ft Menlo Park facility where LLMs predict new materials, robots synthesize them, X-ray diffraction machines validate the predicted properties, and the verification data feeds back into training. “Throwing more compute at the problem is probably having super exponential gains right now per iteration… There’s no saturation in superconductor discovery at all. The bitter lesson is holding, well and alive.”
  • The origin story: a year ago, amid “AI for science” hype, he benchmarked Claude and Gemini on physics and chemistry — “surprise, they sucked.” The missing ingredient wasn’t algorithms but data locked up in national labs, academic labs and semiconductor fabs, absent from internet pre-training. The physical-lab recipe, he says, applies to any domain you want to advance.

2. Four bottlenecks — and culture quietly solves the algorithm problem

  • His full list: context feedback, compute, capital, culture — “and I think culture actually might be the most important bottleneck of all time.” Algorithmic innovation isn’t on the list anymore: with a mission-driven culture, researchers stop being tied to transformers-vs-diffusion tribalism and “the algorithmic stuff takes care of itself.” Two-three years ago architecture was a huge bottleneck; in his view, no longer.
  • Context feedback is step one and where the money is: “context is not necessarily the moat… yet” — VCs are “very quick to analyze moats” — but unique, differentiated feedback loops are where progress is most legible and where the superior business model sits.
  • On superintelligence he’s deliberately deflationary: within-distribution capabilities are genuinely superhuman (“fine, let’s call that super intelligence”), and recursive self-improvement in coding is “totally happening” — but you can’t tell a coding model to “bootstrap a physical R&D lab for me in Menlo Park.” Harry contrasts Cursor’s cloudification with Periodic’s physical data; Midha responds that context is not necessarily the moat yet, while unique feedback loops create legible progress and business advantage.

3. The Cloud Act cracks hyperscaler dominance — the whole Mistral thesis

  • The mechanism: the US Cloud Act gives the US government access to data on infrastructure managed by American companies. So if you’re ASML, or CMA CGM running mission-critical logistics, you legally cannot process that context on AWS, GCP or Azure — and there were almost no trusted European alternatives. Hence Arthur Mensch, “a 33-year-old scientist,” on stage at Vivatech in July 2025 with Macron and Jensen, unveiling a gigawatt facility in Paris. “It’s the first time in 15 years that hyperscaler dominance is up for grabs for startups.”
  • His Mistral thesis, verbatim: “Independence at scale of every part of the AI infrastructure stack” — sovereign land, power and shell, local compute, locally trained and fully open models.
  • Harry’s pushback — don’t Anthropic and OpenAI just come for Europe anyway? On Anthropic, Midha’s answer is that the company has always been “very American aligned,” and the world’s largest enterprise customers — governments and Fortune 500s — increasingly need workloads run locally.

4. Anthropic’s seed: 21 of 22 no’s, and not a single VC firm

  • He’d known Dario “forever” (a lead author on GPT-3); Tom called saying they were leaving to start a new lab. From early 2021 they ran weekly sessions turning “a research hypothesis — scale the scaling recipe — into a business hypothesis”: the AI pair programmer, using the local repository as the context feedback loop, with inference supplying both revenue to buy compute and feedback to improve capabilities.
  • They originally tried to raise $500M and re-anchored to a $100M seed — tiny next to OpenAI’s ~$1B. He made 22 introductions up and down Sand Hill Road and got 21 no’s. The exchange he can’t forget: “Proof? These are the guys who invented GPT-3. How much more proof do you want?” — “What’s GPT-3?”
  • Who got it: ML people overlapping the effective-altruist community — SBF — and Amazon, for whom state-of-the-art models hosted on AWS were “super accretive,” producing the original $4B compute-and-capital-for-equity partnership. Midha invested his life savings (“most of my net worth was tied up in Discord stock”).
  • The kicker comes later, when Harry defends VC as wealth distribution via pensions and endowments: “But how many venture capital firms were in the seed round of Anthropic?” — “None.” — “That’s the answer for you… There’s a huge misallocation of public capital into venture managers.” He has since invested many hundreds of millions across rounds and intends to give most of it away to public benefit causes.

5. Public benefit governance as anti-Congress insurance

  • To Harry’s claim that Congress hearings are inevitable at scale, Midha points to REI and Ben & Jerry’s — public benefit corporations making billions that have never been hauled up — “because they self-moderated.” When held in balance, long-term mission and profit “are not in conflict”; PBC governance is the mechanism for moderating between them, and “we need more public benefit charters in Silicon Valley.”
  • Harry relays a mutual friend’s jab — will these PBC founders not just win their market first? The retort: “Tell them to give me a call when they’d like to be investors in the world’s fastest-growing business of all time. And then they can lecture me about public benefit governance.”
  • The concrete shareholder-illegible decision at Amp: giving away most of its compute at cost — billions of dollars of infrastructure — because the teams truly pushing the frontier “can’t afford price gouging,” and Amp’s stated mission is “to maximize the world’s frontier output.”

6. The Amp Grid: an independent system operator for compute, circa 1885

  • Amp is “not a cloud provider, we don’t own our own data centers, we’re not a traditional venture capital firm” — it’s an independent system operator coordinating capacity so the best teams provision for base load, not peak. His framing: “We are roughly in 1885 industrial revolution England” — frontier labs are factories each “running their own generators in their backyards at half capacity.” Pool the generators so the shoe factory spikes by day and the steel factory by night.
  • Harry’s sharpest pushback: isn’t at-cost compute a loss leader for the venture fund — compute in exchange for a $300M allocation? “That’s not at all… the deal.” He incubates one company at a time: at Periodic he’s on site 3 days a week and has run an 8:00–8:30am daily standup with Liam Donohue for a year, while Amp’s compute team procures upstairs.
  • The numbers: ~1.3 gigawatts secured as proof of concept — roughly $40 of cloud spend over 4 years, financed 20% equity ($10) / remaining debt. Fundraising “is not a problem. It’s actually a systems design problem” that took him a year (really four) to structure. The supply edge traces to a16z’s oxygen program: “step one is you get there first before people realize how valuable it is.”
  • Europe’s sizing, done in gigawatts: Google is “roughly at 12 to 15 gigawatts” that he’s aware of, with a huge land-power-shell pipeline coming. “If Europe does not have access to Google-level infrastructure, then what are you guys even doing?” That’s the bar for full sovereignty over the next 4 years.

7. “Not an AI bubble — a GPU wastage bubble”

  • His refrain: “We are not in AI crisis. We’re not in AI bubble, for sure… We are definitely in a GPU wastage bubble” — billions of dollars of stranded compute sitting unutilized. The cause is that compute isn’t fungible: even within Nvidia, H100s, GB200s and GB300s are “completely different chip types”; a training run on H100s can’t continue on Blackwell, and old clusters are memory-bound for frontier work. “I wish flops were fungible, but not all flops are born equal today.”
  • We’re in compute’s pre-standardization era — electricity in 1885, steel, railroads — where “wars are fought, companies backstab each other.” The exit is what AC/DC did for electricity and TCP/IP’s RFC process did for the internet: technologists propose open standards, bodies like NIST help standardize them. He hopes the industry can “self-standardize… and skip the boom and bust cycles.”
  • The deepest bottleneck is the episode’s opening line: “AI alignment, don’t get me wrong, is hard, but not the hardest problem. Human alignment is really the problem right now.” The unresolved core debate: should a statistical model be regulated and procured like deterministic software (a spreadsheet)? He credits Trump with “trying to do his best… giving America enough freedom to innovate that these standards can even be discovered.”

8. China’s full-stack race — and the case for an inference iron dome

  • China realized “the AI scaling race is not a chip race. It’s a full-stack systems co-design race.” Lacking leading-edge chips, they co-design Huawei silicon with infrastructure and training runs, then run adversarial distillation at scale — distilling Western state-of-the-art from various endpoints, releasing the results as open models, harvesting feedback, iterating until caught up. Then: “Why do we need to open-source anymore?” His verdict, uncomfortable as it is: “It’s beautiful… It’s extraordinary” — the Google integration strategy (land-power-shell, TPUs, Borg, Gemini) replicated with open source as the bootstrap.
  • Does it concern him? “Are you kidding? Absolutely.” Today’s defense is literally his group chats — a founder texts “is anyone else noticing a huge spike in distillation from this region?” and he coordinates informally across his seven boards. The fix is that mechanism scaled: an iron dome for inference, all frontier inference served through a shared proxy that flags attacks across companies. “If we don’t secure frontier model inference… behind a coordinated iron dome, I don’t think we have a sustainable shot at staying at the frontier over the next decade.”
  • What we should know but don’t: insider threats are real, distillation is exploiting US–European disunity and “our political systems,” and mission-critical data centers serving enterprise workloads are “quite vulnerable.”

9. Optimal competition: Thiel was “insufficiently precise”

  • His amendment to “competition is for losers”: perfect competition is for losers — 50 companies training LLMs is restaurants, no defensibility — but monopolies are equally bad: “Monopolies are mafias. Once you have a monopoly… they stop innovating,” hoard resources, and acquire their way up and down the stack — behavior he says he’s already seeing. The target is optimal competition: three or four teams per frontier, good enough for extraordinary returns, uncomfortable enough to keep innovating.
  • Applied to inference: demand can grow combinatorially if bottlenecks unblock — and “if there’s any reason why OpenAI and Anthropic, Gemini and so on don’t hit their revenue targets over the next few years, it’s because they won’t have access to enough compute.” Fifty VC-subsidized inference companies are “lighting hundreds of millions of dollars on fire” in a race to the bottom; the 4–5 winners will be separated by “Supply. Access to supply… If you’re making a steam engine, you need coal.”
  • On his former partner’s tweet (likely Martin Casado) that model creators will reserve the best models: general products amortize development across the most users, like the iPhone, so general models stay broadly available; specialized models get enterprise segmentation. “There’s no one large god model,” and the open-vs-closed access debate is “somewhat overblown.”
  • The category error he’s fought for 4 years: these were never foundation model companies but frontier systems companies — hence Claude Code, hence Mistral Compute. “That was the plan all along… you just weren’t paying attention and you had your neat market maps that your associates were giving you.” Many $100B+ frontier systems companies are yet to be built.

10. Back-to-the-future venture — and “he was right”

  • Frontier industries were founded by investor-cofounders: Arthur Rock at Intel literally wrote the stock incentive plan and ran weekly all-hands; Genentech was incubated in Kleiner’s basement (Herb Boyer plus Kleiner associate Bob Swanson); Mike Markkula was “effectively the first CEO for the first year of Apple.” Midha apprenticed in that mode at Kleiner at 20 under Brook Byers — and thinks check-writing and co-founding are “very hard to coexist inside of one person,” sometimes even one firm. His broader investor method is to treat the future as undetermined, form bottleneck hypotheses, run parallel experiments, and stay willing to be wrong and honest with LPs.
  • His LP advice, quick-fire: “Do the readings. Don’t skip the hard work” — too many LPs outsource capital allocation — and “I would be investing in the bottlenecks.” He fully endorses Harry’s complaint about GPs who’ve never built with AI; his Stanford CS 153 class project is “the one-person frontier lab,” because what took 50 people 4 years ago now takes one. A sovereign fund is running 26 ministers through his year-long program — no graduation certificate until each builds and deploys an agent.
  • What makes Dario special from inside: “sheer scientific brilliance,” but more specifically he’s “a physicist at heart… an empiricist,” plus mission clarity — “No drift. We won’t take shortcuts… willing to make huge trade-offs” — which attracts world-class talent despite “crazy” ad hominem attacks.
  • The change of mind: health experiences in his family and himself made him “stop taking for granted” time — his first “life scaling law” for students. And the tombstone answer, blurted at a San Francisco dinner party when his wife Viv put him on the spot: “He was right. And the room just went dead quiet and they were like, yep.”