[State of Evals] LMArena's $1.7B Vision — Anastasios Angelopoulos, LMArena
Summary
Arena’s $100 million raise buys strategic retries, not a mandate to burn. Anastasios Angelopoulos calls capital “cards to flip” if the first bet fails; immediate costs include funding free inference, hiring, and replacing Gradio with React, but he stresses that Arena need not spend the entire raise.
Arena’s moat is the scale and realism of organic usage, not a static benchmark catalog. The host described the community as 5 million MAU; Angelopoulos cited roughly 250 million conversations on the platform and mid-tens of millions monthly. He also said roughly 25% of users do software for a living. Unlike arenas built around pre-generated outputs, Arena captures users asking their own questions, continuously refreshing the evaluation distribution.
The public leaderboard is a credibility-building loss leader with an explicit no-pay-to-play covenant. Angelopoulos calls it both a “charity” and a “loss leader”: released models appear regardless of payment or score, and providers cannot pay for removal. His answer to the “Leaderboard Illusion” critique is that its analysis contained factual errors—including a claimed 9% open-source sampling rate versus Arena’s roughly 60/40 mix—and mischaracterized long-running preview testing.
Both guest and host reversed their skepticism on image generation after Nano Banana demonstrated its economic pull. Angelopoulos now expects multimodal systems to become among AI’s most economically valuable consumer and enterprise capabilities, with marketing and design among the fastest-growing adoption segments. The host’s specimen was feeding DeepSeek V3.2 explanations of RL environments into Nano Banana Pro and receiving a paper-quality diagram that might once have taken a PhD student a month.
Arena is expanding from one aggregate ranking into occupational, multimodal, and agent-specific evaluation categories. Single-digit shares of its large audience already represent medicine, legal, finance, accounting, creative, and marketing cohorts; video is planned for later in the year or early the next. Code Arena could also evolve from evaluating models toward comparing full harnesses such as Devin.
Consumer retention and startup focus remain the execution constraints. Persistent history made sign-in a meaningful retention driver, but Angelopoulos says “every user is earned” and can leave at any moment after a “lightning-in-a-bottle” spike. An API remains possible, though Arena’s current answer to strategic sprawl is simple: “We really should be doing one thing well”—arenas.
Deep dive
1. Company formation turned an academic benchmark into infrastructure
Angelopoulos’s deciding premise was that Arena could not reach the necessary distribution, platform quality, or operating scale as an academic project or nonprofit. The mission was already clear: “measure, understand and advance frontier AI capabilities” through real users and organic feedback; the company became the practical scaling structure.
Anjney incubated Arena, providing early grants and resources before the founders had committed to a business; Sequoia also provided a grant. Anjney formed an entity with the unusual assurance that the team could walk away. Angelopoulos ultimately agreed that forming a company was “the only way that we could scale.”
The $100 million raise is optionality: “The purpose of money at a company is to give you cards to flip.” Arena funds all inference for free platform usage, hires headcount, and migrated from Gradio—which carried it to roughly 1 million MAU—to React for richer components and a developer pool more familiar with that stack.
2. Organic prompts are Arena’s core data advantage
The host framed Arena’s community as 5 million MAU. Angelopoulos separately cited “more than 5 million” without specifying the metric, about 250 million conversations on the platform, and mid-tens of millions of conversations each month. He said roughly 25% of users do software for a living, while approximately half now log in, giving Arena more ability to study behavior alongside surveys—though he explicitly preserves the caveat of response bias.
The host relayed Artificial Analysis’s “Gartner of AI” ambition. That group aggregates and independently reruns public benchmarks into analytics and reports, whereas Arena lets people enter their own use cases and questions rather than merely judge pre-generated outputs.
The host supplied the counterpoint that curated examples can teach weak prompters what is possible. Angelopoulos agreed that viewing other people’s prompts is educational, while distinguishing Arena’s own-use-case input as the source of its realism.
3. Leaderboard integrity is the asset Arena refuses to sell
The “Leaderboard Illusion” paper alleged that preview-model testing created undisclosed inequities. Angelopoulos called the critique “unscientific,” pointing to some corrected factual errors and, specifically, its claim of roughly 9% open-source sampling against what he says was closer to a 60/40 distribution.
His defense of previews is cultural as well as statistical: Arena has long exposed prerelease models under secret code names, and its community enjoys discovering them. Nano Banana began there and became a “global sensation.”
Angelopoulos said Nano Banana’s moment alone changed Google’s market share. The host argued that it was visibly ahead of everything else and linked it to billions of dollars moving in Google’s stock.
The host’s sharper objection was that not every preview reaches the leaderboard. Angelopoulos’s answer: unreleased models need not appear, but every released model receives a statistically sound score derived from millions of votes. Providers cannot buy inclusion or pay for removal—the ranking must remain a “transparent and fair reflection of model performance.”
4. Multimodal models changed both speakers’ economic forecasts
The host admitted he once viewed image generation as peripheral to AGI and reputationally troublesome: why not concentrate AI’s upside on language, coding, and reasoning? Nano Banana changed his view; Angelopoulos likewise said, “I was also kind of wrong about this.”
Their revised thesis is practical rather than philosophical. Marketing and design are among the fastest-growing AI adoption segments, while creators gain an “infinite supply” of diagrams, explainers, and infographics—making multimodal systems potentially among AI’s most economically valuable consumer and enterprise capabilities.
The best specimen was DeepSeek V3.2: the host fed its RL-environment explanations into Nano Banana Pro and generated a diagram that helped him understand the paper. Producing comparable work manually, he argued, might once have taken a PhD student “like a month.”
5. Arena’s roadmap broadens evaluation without abandoning focus
Angelopoulos wants Arena to remain the industry’s “north star”: a continuously fresh benchmark that resists overfitting because new data points constantly enter. Arena tracks new models and use cases and has released millions of real conversations so researchers can study and improve real-world performance.
Occupational views now expose results for medicine, legal, business, finance, accounting, creative, and marketing users. Even single-digit percentages become meaningful cohorts at Arena’s scale; video evaluation is expected later in the year or early the next.
An API is possible, but Angelopoulos questions it on startup-focus grounds: a company should do “one thing well.” On retention, persistent history increased sign-ins, but no feature removes the basic obligation that “every user is earned” each day.
He is also looking for experts in consumer product, machine learning, B2B go-to-market, marketing, and related areas to join Arena.
Code Arena may extend the unit of evaluation beyond a model to the full agent harness. In a potential Cognition partnership, Angelopoulos proposed putting Devin into Arena and testing whether it is “the best, or one of the best in the world, at doing what it does,” in response to people saying Devin was dead.