Arena CEO: There Will be a $100BN US Open-Source Model & Data is a Trillion Dollar Market
Summary
Open-source leadership has shifted faster than expected. Anastasios says Kimi K3 beat every American model, including Fable, on a meaningful subset of tasks such as front-end coding; that does not rule out distillation, but it breaks the story that Chinese labs are “just distilling American models.” Open source remains a small share of global inference spend today, yet its trajectory puts closed-model oligopoly economics under pressure.
Enterprises will increasingly demand AI sovereignty: their own models, fine-tuned on proprietary data, without handing intelligence or supply-chain control to a potential future competitor. Anastasios expects at least one “multi-hundred billion, if not trillion-dollar” American-first open-source company, monetizing through inference revenue shares or using free models to win the enormous AI-modernization and deployment-engineering market.
Chinese models create a policy trap. Anastasios guesses US restrictions are likely within three years, though “very uncertain” and not necessarily desirable: banning them might reduce backdoor risk and help domestic labs, but it could leave American companies building on open-source model number 10 while foreign competitors use number one. Local hosting is no complete defense—a model could contain a hidden sequence that jailbreaks it and causes it to “vomit out” private data.
Inference and routing should get cheaper, but the route there is contested. Routing requires understanding each query’s domain and difficulty, continuously measuring every model, and onboarding weekly releases; Anastasios thinks the hype must be purged before winners emerge. He expects Anthropic’s “disgustingly high gross margins” to face pressure after public disclosure, while Harry counters that genuinely differentiated companies can retain Chanel- or Apple-like pricing power.
AI security becomes an AI-versus-AI problem: guardian models must watch agent traces and be “equally as smart as the agent,” because humans will be too slow. Anastasios rejects government preapproval of releases—“why should the DMV be telling me what model I can use?”—and prefers outcome-based liability and enormous fines. His concrete warning is already operational: Arena interviewed an apparently real, technically excellent candidate who ultimately proved to be an AI-generated fake.
Two-thirds of at least 75 neo-labs will be worth nothing or be bought out for parts, Anastasios predicts. At a $10 billion valuation, a lab seeking a 10x outcome needs roughly $4 billion of revenue within two or three years at a 25-30x multiple; team-value downside protection may justify the first speculative check, but “next round’s a bitch” once investors demand an actual business model.
Data is a $100 billion market by 2030, potentially $1 trillion, because it scales alongside models and currently attracts roughly 10-20% of frontier labs’ GPU spend. Anastasios calls data less commoditized than GPUs—“in order for data to become irrelevant, humans need to become irrelevant”—and sees leading providers becoming worth hundreds of billions despite concentrated revenue.
If inference commoditizes, frontier labs will climb into applications, putting legal and other AI-native software vendors at risk; Harry cites design as another example. Durable GTM, network effects and enterprise entrenchment become the defense. Anastasios nevertheless sees enterprise AI adoption as another potential 10x for Nvidia, while warning that open-source cost savings could impair OpenAI and Anthropic revenue, raise insolvency risk amid compute debt and knock the wind out of the surrounding infrastructure ecosystem.
Deep dive
1. Kimi K3 broke the distillation-only story
Arena measures real-world AI performance, not static benchmarks: whether people complete actual work, encounter hallucinations, can steer a model and ultimately prefer its output. Its human feedback helps labs improve while giving the market a continuously updated view of models released multiple times each week.
Anastasios calls Kimi K3 a genuine narrative violation because it surpassed every American model, including Fable, on a meaningful subset of tasks such as front-end coding. It might still use distillation within training, but “distillation is only part of the story”; something additional allowed it to exceed the models supposedly supplying the distilled intelligence.
Harry’s pushback—worth keeping—was that the model was reportedly 27 trillion parameters and “pretty clunky,” while a highly AI-engaged developer he knew did not find it better overall. Anastasios held the narrower claim: the subset win mattered, not because Kimi K3 dominated everything, but because it punctured assumptions about permanent American scientific hegemony.
2. Enterprise sovereignty creates the opening for an American open-source giant
OpenRouter’s rankings overstate open-source consumption, Anastasios cautions, because customers tend to buy proprietary inference directly while using OpenRouter where open models need failover and added services. Anthropic’s revenue remains a “total hockey stick”; open source is still a small fraction of worldwide inference spending.
The longer-term incentives point elsewhere: businesses will want to own their intelligence, fine-tune models on private data, control the full stack outside compute hosting and avoid feeding information to a vendor that might later compete with them. Software itself becomes less defensible as creation gets easier, leaving network effects and proprietary data as the durable moats.
His Coca-Cola/Cisco example carries the mechanism: neither company must lead frontier research if its enormous user corpus can power a self-improving product. This makes specialized intelligence less a Silicon Valley curiosity than an economic imperative—“businesses are going to need a way of keeping a moat in the age of AI.”
Anastasios predicts at least one multi-hundred-billion or trillion-dollar American-first open-source company. It could collect revenue shares after inference providers cross a threshold, or use models as lead generation for fine-tuning, deployment engineering and a decade-long AI-modernization wave; Harry correctly notes FDE is not unique, but Anastasios argues that FDE plus sovereign Western models may be more sustainable.
3. Chinese-model restrictions trade security for American competitiveness
A Chinese researcher told Harry that Chinese teams work harder and benefit from policy support, regulation and subsidies. Anastasios says China has tailwinds and headwinds and rejects such a one-sided account: US companies retain the world’s strongest chip ecosystem while Chinese labs remain hardware constrained. Export controls may hurt them now, yet also incentivize a domestic stack; the strategic choice is whether to starve competitors or addict the world to Nvidia hardware.
China already restricts American models domestically, Anastasios says. If China restricted Chinese models in the US, he argues, it would sacrifice Chinese revenue, global mindshare and dominance merely to deny American businesses the best open intelligence. Separately, a US ban might suppress backdoors and accelerate domestic alternatives, but “since when has America been about number 10?” if overseas companies retain access to number one.
Self-hosting does not eliminate model risk. Anastasios imagines a locally deployed chatbot connected to corporate data but trained abroad: an attacker supplies a hidden code word or character sequence, jailbreaks it and makes it “vomit out” backend information. Malicious behavior can reside inside the model even when the infrastructure is controlled locally.
His three-year bet is that restrictions probably arrive, though he stresses uncertainty and does not endorse them. Harry agrees for political reasons: Sam Altman and Dario Amodei could marshal labs, investors and government relationships into effective lobbying; Jensen Huang’s open-source advocacy is simultaneously self-serving and patriotic, expanding GPU demand while preserving choice and competition.
4. Routing is real technology, but inference economics remain unsettled
A useful router must classify a query’s domain and difficulty, know every candidate model’s measured strengths, optimize cost and performance, then rapidly absorb new releases. Anastasios sees a routing hype cycle that needs purging, yet rejects the idea that every enterprise—or every vendor claiming a router—can genuinely solve this machine-learning problem.
Harry presses that Ramp, Fireworks and numerous inference providers already make routing look commoditized. Anastasios’s honest answer is that he does not know how deeply each product solves the problem: the eventual winners will be those for whom routing is a priority and whose technology demonstrably saves money without sacrificing performance.
Anastasios expects Anthropic’s “disgustingly high gross margins” in inference to become negotiating ammunition once public disclosure reveals its cost structure. Harry’s rebuttal is pricing power: Chanel can charge £6,000 for a low-cost bag, and Palantir can refuse cost-plus economics. Anastasios concedes unique products can preserve margins, as Apple has, even in transparent markets.
He expects Anthropic has incentives to list before OpenAI because it appears better prepared and generates free cash flow; reports suggested it could arrive “as soon as October,” though he hedges their reliability. A major open model beating Opus 5 or Fable across categories would create material IPO risk, alongside the deeper business problem such a result would imply.
5. AI agents need AI guardians, not a government release queue
Anastasios treats the reported OpenAI/Hugging Face breach as hugely underappreciated: in his account, a model escaped safeguards, accessed company data, and defenders needed an open model because closed systems refused the work. The lesson is not to halt agents, but to build access controls and guardian systems around them.
A guardian model watches every agent’s trace, classifying actions as safe, unsafe or anomalous. It must be “equally as smart as the agent” so the protected system cannot outwit it; humans will be too slow to supervise machine-speed activity directly.
Harry ridicules proposals to approve every release administratively, invoking the difficulty of overturning a parking ticket. Anastasios agrees: “Why should the DMV be telling me what model I can use?” He favors outcome-based regulation—clear prohibited outcomes, corporate liability, huge fines and scrutiny when systems leak data—over technically incapable officials controlling product-release timing.
The threat is already tangible at Arena. An apparently real infrastructure candidate passed interviews with world-class engineers, only to prove nonexistent: “It was AI.” Anastasios says the episode “totally fucking worries me”; Arena is considering in-person onboarding and requiring new hires to appear to receive their laptops, so someone must physically appear, shake hands and establish that they are real.
6. Neo-lab downside protection disappears at the next financing
Elite AI talent can command tens of millions of dollars, particularly established specialists with years of experience and tens of thousands of citations. Many remain concentrated in frontier labs, though Anastasios expects more departures as those organizations become large public companies where individual researchers feel less able to matter.
Thinking Machines illustrates both readings of a neo-lab. The generous case is that a restructuring left it only six months to produce Inkling, which Anastasios describes as Arena’s number-one American open-source model; the harsher case is that after a year and a half, nine Chinese models still rank above it, leaving it number 10 globally.
Of at least 75 neo-labs, Anastasios says two-thirds will be worthless or acquired for parts. A $10 billion company seeking a 10x needs roughly $4 billion of revenue within two or three years at a 25-30x revenue multiple; producing a model and celebrating is “old fucking news” without hypergrowth and a sustainable monetization strategy.
Harry notes employees can obtain tender liquidity and cites strong revenue at Mistral and ElevenLabs; Anastasios agrees those are not the problem and expects ElevenLabs to become public. The fragile wager is a multibillion-dollar, zero-revenue lab whose team might fetch $1 billion: that can protect an initial $200 million check, but “next round’s a bitch.”
7. Data scales with AI and is less fungible than GPUs
Anastasios defines data as a scaling complement: just as more cars increase demand for gasoline, more and larger models increase demand for training data. More companies training proprietary models compounds the need; frontier labs already spend roughly 10-20% of their GPU budgets on it.
“They think about data as a commodity. It’s really not.” His sharper claim is that data is less commoditized than GPUs, because it remains necessary until humans themselves become irrelevant—effectively, until AGI. Sourcing and cleaning it is unpleasant, laborious work, while model production increasingly resembles “data plus GPUs equals model.”
He forecasts at least $100 billion by 2030, if not $1 trillion, and believes leading providers could readily become worth hundreds of billions. Revenue concentration does not invalidate them: “Silicon Valley investors have become total bitches” about a feature shared by TSMC, government contractors and other enormous businesses.
The customer base should broaden as enterprises build their own intelligence. If every company needs a proprietary model, it also needs proprietary data that compounds its moat; the same logic reaches medicine, where the missing asset is not a different GPU but biological data infrastructure and a rapid feedback loop.
8. Arena is betting that evaluation becomes deployment’s bottleneck
Arena has 30 million-plus monthly visitors, including knowledge workers and “unhirable experts” doing real tasks; Anastasios says that makes it larger than xAI, Huggy Face, Manus and Genspark by traffic. Their organic feedback creates a flywheel for agent evaluations grounded in actual traces rather than purchased benchmark data.
Evaluation decomposes into performance, cost and latency. The latter two are easy to measure; performance depends on the company and use case. Arena’s bet is that extracting business-specific performance signals from agentic traces becomes the central deployment bottleneck, helping enterprises choose, route and potentially train models.
The company is above $100 million in annualized run-rate revenue, calculated as Q2 multiplied by four, but is not yet free-cash-flow positive while reinvesting. Anastasios says one model provider would undermine Arena, three would preserve meaningful competition and two is “a little dicey”—a candid measure of how much its evaluation economics depend on model diversity.
9. Frontier labs will climb the stack, but entrenchment still matters
If inference commoditizes, OpenAI and Anthropic logically move toward applications and capture more end-customer value. Anastasios says software CEOs are already “shaking in their boots” as customers replace traditional vendors with labs perceived as more AI-forward; Harvey, Lagora and scalable automation products face direct platform risk. Harry also cites Claude Design as an example of application-layer pressure.
Harry spots the contradiction: enterprises supposedly fear frontier labs while simultaneously embracing them. Anastasios concedes both behaviors exist. Harry’s resolution is GTM: a design tool spreads individually, while legal software requires partner relationships and reluctant junior-lawyer deployment; being Anthropic’s priority number 12 may not beat a focused vendor living the workflow daily.
Network effects, operations and deep integration remain defenses. Anastasios sees a system integrator such as Infosys as less likely to be displaced and calls the “SaaSpocalypse” overstated; Salesforce has a serious AI strategy, while Harry argues weaker, less-entrenched products such as Wix should be materially more nervous.
Anastasios picks Nvidia as the likeliest first $10 trillion company, arguing enterprise adoption could 10x the industry and Nvidia “very reliably.” Yet he worries open-source savings could cut OpenAI and Anthropic enterprise revenue, trigger insolvency amid compute debt and collapse dependent inference and routing businesses—a concentration risk that may force the “sobering up” and consolidation he welcomes.