Inside the Race to Measure Frontier Intelligence
Inside the Race to Measure Frontier Intelligence
Summary
- The founding thesis of independent AI evaluation got its proof point when Meta’s Llama 4 “performed worse” in closed private tests while showing “incredible ability” on major public tests — a gap between claimed and actual capability that labs’ self-reports can’t close. Rayan Krishnan argues labs investing billions need a “rational purchasing market” with third-party evidence, a call he says Demis and others across the industry are echoing. “Every time a new trillion-dollar industry emerges, there is a need for such an independent testing group.”
- The firm’s core governance choice is refusing to sell training data to labs, invoking Enron as the cautionary tale. Krishnan says “a significant portion of this industry has created performance benchmarks as a mechanism to sell its data,” and when auditor and consultant are the same party, “the main thing becomes to pass the audit or, in this case, beat the benchmark” — a structural reason for avoiding conflicted incentives.
- Token spend could begin to rival or dwarf payroll, and enterprises are rationing intelligence arbitrarily. One Fortune 10 company gave engineers a roughly $100/day Cloud Code budget with limits resetting at 4 PM — so “the most productive working hours” fell between 4:00 PM and 6:00 PM. Separately, another Fortune 10 company raised an arbitrary per-employee limit from $100 to $300. In Jennifer Li’s unlimited-usage experiment, engineers burned 1–2 billion tokens a day (one hit 6 billion), totaling about $1.5 million in a month — “10 times more on tokens than we did on employee salaries.”
- Model selection is genuinely counterintuitive without private-repository testing: “in many cases Sonnet is more expensive than Opus because it consumes too many tokens.” The firm’s Val Smith product builds internal benchmarks from a company’s GitHub codebase to find the Pareto-optimal agent; Li’s team’s audit found Cognition’s Devin unusually token-efficient and that subscriptions can beat per-token pricing. Krishnan’s frame is that “a firm is effectively its evals” — companies that make eval results legible to ROI math will outcompete.
- The frontier benchmark is the Recursive Self-Improvement Index — proxy indicators across pretraining, post-training, and framework-level engineering, since actually letting a model train its successor is “very expensive and slow.” Krishnan calls RSI a major long-run concern: “one country or one company leaping forward and creating models that we know little about,” which is why the discussion cites researchers calling for joint government negotiations.
- Ben Horowitz’s policy division of labor: government defines what it fears and enforces rules; private evaluators answer the two operative questions — “can the model do it, and can the model be made to do it?” Government is “not particularly equipped” to evaluate, “but they are very good at setting rules because they can enforce them.” Rayan adds that labs’ posture of keeping dangerous models in-house “doesn’t always work.”
- On geopolitics, Krishnan wants a “common language of assessments” playing the arms-control verification role — Reagan’s “trust, but verify,” with eval frameworks as the overflights. He admits surprise at “extremely inefficient” sovereign-AI duplication, notes a Xi–Trump meeting next month as a sign of trust without any verification mechanism, and says cyber evals must move beyond code vulnerabilities to modeling corporate cloud and energy-grid environments.
Deep dive
1. Llama 4 exposed the self-report problem — and the Enron lesson shapes the business model
- Krishnan’s origin story: in early 2024 his team concluded “all the public tests were not enough to measure the progress of the models.” The proof came with Llama 4 — “a bit of a disaster” — which “in our closed, private tests… actually performed worse” while demonstrating “incredible ability” on major public tests where questions and criteria are open. Labs want a “rational purchasing market” where billion-dollar investments point to real evidence, “not just self-reports to justify the investment.”
- Ben’s analogy for fuzzy capability standards: the MPAA. What’s R versus X “has changed over time,” and there’s no answer beyond “I’ll know it when I see it” — but “if enough people who run companies, manage finances, or whatever else, agree, then it will become the norm,” which beats today’s open benchmarks that are “vulnerable to attack, and also too narrow.”
- Krishnan’s structural commitment, drawn from auditing’s failures: never sell training data to labs, despite being “often pushed to do this.” His Enron read — when the same group audits and consults, “the main thing becomes to pass the audit or, in this case, beat the benchmark. And that’s not at all what the market benefits from.”
2. The eval machine: six-hour windows, Steve, and “there’s always a higher peak”
- The operating constraint: “we never want to be a delay or a hindrance to the release of a model.” What began as Krishnan and his co-founder Lynx working all night pre-launch is now massively distributed infrastructure running at each model’s rate limits, plus an internal system called Steve intended to progressively absorb more human work.
- The Recursive Self-Improvement Index, the benchmark he’s proudest of: since having a leading model train its successor is “very expensive and slow,” the firm builds proxy indicators for each stage — pretraining, post-training, and framework-level engineering — while examining the mechanisms and behaviors that enable high-quality research and creation. Labs discuss RSI in model data sheets, “but there is still no common language.”
- On benchmark retirement, the unofficial T-shirt slogan is: “There’s always a higher peak.” As labs hill-climb, “our job is to constantly build these new mountains for them.” Torenberg also argues that benchmarks must track the state of the world, such as updated case law in legal-research evaluations, much as lawyers retake licensing exams.
- The agentic shift changes eval architecture: from ImageNet’s millions of one-to-one labels to tasks like “create 50 fully functional web applications” — smaller sample sizes, far richer rubrics, and infrastructure stable enough to resume mid-trajectory on tasks running “hours, days, sometimes weeks.”
3. Tokens versus payroll: the enterprise is flying blind
- Krishnan’s telling anecdote: one Fortune 10 firm deployed Cloud Code with a budget of about $100 a day for its engineers, with a request limit resetting at 4 PM — creating a dead period during the day and peak productivity from 4–6 PM. Separately, another Fortune 10 company arbitrarily set a $100-per-employee limit and later raised it to $300, “almost like one employee’s salary spent on tokens.” His diagnosis: “misjudgment of intelligence at every level of the stack,” compounded by Anthropic’s large model-maintenance costs and low margins.
- Jennifer Li’s own experiment: a month of unlimited coding tools saw engineers spending 1–2 billion tokens daily, one peaking at 6 billion — roughly $1.5 million in tokens, “10 times more on tokens than we did on employee salaries.” Her team’s audit of logs, traces, and GitHub work yielded “pretty strange insights”: Devin from Cognition “uses tokens very effectively,” and subscription pricing can beat per-token pricing.
- That work led to the product Val Smith: build internal benchmarks from your own GitHub codebase to find the Pareto-optimal agent. Results are counterintuitive — “in many cases Sonnet is more expensive than Opus because it consumes too many tokens” — amid a confusing middle tier that includes Luna and Terra alongside Opus and Sonnet, plus Spark, whose version 1.2 is described as very functional. On Torenberg’s question about real-time routing via OpenRouter (“which Stripe just bought,” as stated), Krishnan says the name is “a bit of a misnomer”: it is primarily a gateway, and “the really hard part of routing is creating evaluation systems.”
- The thesis line: “a firm is effectively its evals.” As Torenberg suggests token spending could “overshadow payroll spending,” Krishnan says the companies that make eval results legible for ROI calculation “will outperform competitors in the long run.”
4. Policy: rules from government, verification from evaluators
- Ben’s division of labor, in full: government has some idea what it fears — biohacking, cyberattacks — but the operative questions are “is the model capable of doing this? And can you make a model do this?” Government is “not particularly equipped” to evaluate that over time — “simply an inappropriate state function” — “but they are very good at setting rules because they can enforce them.”
- Rayan adds that labs’ promises to keep dangerous models in-house do not always work: “even that doesn’t always work… we live in interesting times.”
- Jennifer’s stance is that policy debate has been “very abstract,” so the firm stays in “evidence-gathering mode,” while regularly briefing the executive and legislative branches. She insists “it’s not really our job” to recommend policy direction. Torenberg separately raises the compute threshold of “10 to the 26th power of FLOPs.”
- On misuse testing, Jennifer places the work within the consistency category: “there are already cases now where models undergoing cybersecurity evaluations are actually hacking bug bounties and finding other ways to get around restrictions.”
5. Geopolitics: evals as the verification layer of an AI arms race
- Krishnan’s honest surprise: from “my very idealistic perspective,” sovereign AI looks “extremely inefficient” — duplicated data centers, data-provisioning processes, and giant training runs — “but it seems that we are not living in such a world.” His proposed response is a shared assessment language for discussing risks and areas of agreement.
- The nuclear-arms analogy: Reagan’s “trust, but verify.” Trust signals exist — “Xi Jinping and Trump are going to meet next month” — “but there is no clear way to actually perform the verification part.” Common assessments would play the role of warhead counts and overflights; RSI is a key long-run concern because a country or company could leap ahead with “models that we know little about.”
- Jennifer’s values point, from lived experience: born and raised in China and a user of Chinese open-source models, “you still can’t let them talk freely about the CCP and the whole history there” — evals inevitably encode the developer’s values, complicating standardization across labs and countries.
- The frontier of the eval catalog itself: cyber work is shifting from code vulnerabilities and memory leaks to infrastructure-level risk — “modeling larger-scale environments of corporate cloud infrastructure or even energy grids.” Closing conviction: the most valuable model for this business is one “where incentives are aligned with conducting quality assessments,” rather than with helping develop intelligence or the methods by which models improve.