METR’s Joel Becker on exponential Time Horizon Evals, Threat Models, and the Limits of AI Productivity
Summary
METR’s headline curve shows AI climbing a remarkably straight capability trend, but its “time horizon” measures human task difficulty—not agent runtime. At 50% reliability, the suite spans tiny SWA actions through 20–30-hour HCAST work and RE-Bench research engineering. The caveat is material: tasks are neatly scoped, often auto-gradable, mostly non-visual, and stripped of tacit organizational context, so the chart is not a proxy for all real-world work.
Opus 4.5 was a large enough jump to change experienced engineers’ behavior and challenge Becker’s preferred seven-month doubling trend. He watched developers go from resisting coding assistants to “practically not writing a line of code”; its result fits the faster four-month curve, while falsifying his seven-month line “in some way.” Still, Becker argues that one release can reflect task-distribution noise: the more informative signal is progress over one to three years.
The original finding that AI slowed developers cannot simply be rerun against today’s workflows. Developers now resist being randomized into AI-disallowed work, creating selection bias, and increasingly work on several issues or workstreams concurrently. Apparent productivity can also be overstated: AI enables projects whose counterfactual speedup is “maybe infinite,” but those projects were often deferred because they were less valuable to Becker—and organizations cannot necessarily absorb 10 times more output.
Becker’s phase-change concern is not impressive benchmark performance alone but a fully closed AI-R&D loop. “Ninety percent automated isn’t enough”: the remaining tail may include failing GPUs, data-center operations, cooling infrastructure, chip design, or chip production. METR’s GPT-5 and GPT-5.1 reports concluded that those models were insufficiently capable for catastrophic harm, yet Becker says full-loop automation could create conditions for a destabilizing capability explosion.
Compute is both the visible accelerator and a plausible brake on capability growth. If algorithmic discovery itself requires expensive experiments, slower compute growth hits twice—less raw scaling and less algorithmic progress—potentially delaying major milestones. That thesis remains conditional: some ideas require little compute, AI labor could accelerate research, and the discussion noted that OpenAI-based figures provide little visibility into spending by Meta, xAI, and DeepMind.
Single leaderboard scores conceal the exact “secret 11th thing” that may determine whether autonomy works. Becker wants more open-ended evaluations, agent transcripts, and tests of whether code would actually merge into main—not merely pass unit tests. The hosts noted that harnesses can move performance by roughly 10 percentage points, making scaffolding valuable today even if “across model generations, it’s not so valuable.”
Becker’s Manifold win illustrates how prediction markets can price agency—and potentially privileged information—rather than pure forecasting skill. He spent roughly $5,000 on charity to move a donation market, won fake currency, and became its most profitable trader; meanwhile, an AI-model market drew $28 million in volume despite insiders potentially knowing the answers. His updated view is guarded: calibrated probabilities have value, but “gambling-like behaviors are socially costly.”
Deep dive
1. METR connects capability measurements to specific catastrophic-risk cases
Becker defined METR as model evaluation plus threat research: measuring what systems “might look like today and tomorrow,” what they will actually do in deployment, and whether those capabilities and propensities connect to “enormous or catastrophic risks to society.”
METR’s GPT-5 report, followed by analogous work on GPT-5.1, concluded that the models did not pose those large-scale risks. Becker’s reasoning was capability-based: despite impressive benchmark scores and everyday usefulness, METR believed they were not capable enough to execute the required harms.
The organization has shifted emphasis away from autonomous replication—an AI acquiring resources and setting itself up independently—and toward R&D acceleration inside a lab. That scenario could produce a capability explosion and become destabilizing.
Swyx highlighted METR’s posture as a separately funded organization rather than a lab-funded evaluator. Becker noted that METR came out of ARC and argued that without an independent source of expertise, he could “bang that drum forever” without improving the information available to society.
2. Time horizon measures task difficulty, not how long an agent stays alive
The curve began as a scattered 2023 internal slide: capability on one axis and time, compute, or other resources on the other. Once METR operationalized capability as human task duration at 50% model reliability, the empirical line became “remarkably straight”—far straighter than the original intuition sketch.
Becker stressed the recurring misreading: a five-hour time horizon does not mean a model works productively for five hours. It means the model can reliably solve tasks that take a human about five hours, even if the model finishes in “zero minutes or five minutes.”
The roughly 170-task distribution progresses from SWA-style atomic actions—the file named passwords.txt probably contains the passwords—through HCAST tasks reaching 20–30 human hours, then RE-Bench’s very challenging novel machine-learning research-engineering challenges.
Task selection limits the claim. METR favors economically valuable, scalable, often auto-gradable tasks that a skilled “low-context human” could solve from supplied information; vision-heavy, open-ended, externally interactive, and tacit-context work is underrepresented. Real work is substantially more “messy” than the benchmark.
3. Opus 4.5 challenged one trend line while reinforcing the longer trend
Becker called Opus 4.5 a meaningful bump both on benchmarks and in practice. Some highly capable engineers he knew moved from being selective or resistant about AI coding to “practically not writing a line of code,” a discontinuous behavioral change he also felt personally.
Opus 4.5 fits the faster four-month time-horizon doubling line proposed when METR published its work, but not Becker’s preferred seven-month line. He conceded that it was “falsifying my trend line in some way,” while remaining unsure whether the deviation reflected latent capability or that release’s fit with METR’s task distribution.
The hosts challenged anecdotal claims that Claude Code had run for five or 30 hours: it might have spent much of that time “doing absolute bullshit,” and one successful run could fail on repetition. Becker’s answer was to treat anecdotes as evidence, but weight one- and three-year trends more heavily than individual releases.
4. AI productivity has outgrown the clean design of the original RCT
METR has been redoing its developer-productivity work, but Becker would not disclose results. Replication is harder because developers who expect meaningful gains are increasingly unwilling to be randomized into “AI disallowed,” leaving researchers with tasks whose owners already suspected AI would add little.
Workflow has also changed since approximately March 2025. Developers now run multiple issues or lines of work concurrently, whereas participants in the earlier study largely supplied all their issues and worked less concurrently; randomizing one isolated task no longer captures how coding actually happens.
Becker separated speed on an unchanged task from expansion into new work. Several of his side projects would not exist without AI, making the apparent speedup “maybe infinite,” yet their value is not infinite: they were less valuable to him, which helps explain why he had not acquired the expertise to do them before.
Alessio added the organizational constraint: even if AWS engineers became 10 times faster, customers could not absorb 50,000 additional services. Becker agreed that optimistic self-estimates are easy to overstate, while emphasizing that people at AI companies are probably still being significantly accelerated.
5. Full-loop R&D automation is Becker’s phase-change threshold
Becker’s current safety judgment mixes formal evaluations with observation. Models still look “kind of derpy” in transcripts, misuse resources, and display obvious faults; slightly weaker systems have been broadly deployed for six months without extraordinary danger, making catastrophe from a modest incremental improvement surprising on prior evidence.
Swyx’s pushback was that emergent capabilities may fuse discontinuously, weakening any “n minus one was fine” argument. Becker agreed continuity is “flimsy” given the small number of frontier releases, but said the surprising regularity so far gives him some faith that progress might remain continuous.
The breakpoint that would truly concern him is fully automated AI R&D within a lab. Even a one-year measured time horizon would remain ambiguous because “ninety percent automated isn’t enough”; some unmeasured final 10% may prevent the feedback loop from closing entirely.
The loop might be software-only, where better models recursively produce still better models with fixed hardware, or it might require chip design and ultimately chip production. Becker could not rule out an explosion once such a loop closed: “Who knows what happens after that point.”
6. Research benchmarks capture only a slice of the automation loop
Paper-reproduction or novel-research benchmarks measure a genuine component of AI R&D, but Becker’s “contrary view” is that they omit a long operational tail. GPUs fail, somebody must repair a data center, and somebody must call the water company when cooling breaks—capabilities not represented by a paper score.
Swyx argued for a public “wagon wheel” of perhaps 10 critical capabilities instead of collapsing everything into one number. Becker agreed dimensionality is lost, but predicted any list would reveal “a secret 11th thing” that was difficult to specify beforehand and obvious only after it became binding.
The host suggested versioning the list annually, as security communities do with top-10 risks. Becker welcomed the challenge, while retaining the deeper uncertainty: METR may currently measure only a small proportion of what full automation requires, so closing the entire loop could arrive later than benchmark extrapolation implies.
7. Compute slowdown could compound through algorithmic discovery
Becker’s model starts by taking the time-horizon trend literally, then asks what might bend it downward. Compute is an obvious candidate: if capability depends directly on compute and algorithmic progress also depends on compute-intensive experimentation, slower growth can reduce both components.
Transformers, RLHF, learning-rate schedules, and other improvements are not merely flashes of labor; researchers often need scale to reveal an algorithm’s gains and “a ton of experiments” to discover it. Under the strong assumption that compute is the bottleneck, halving compute growth could roughly halve time-horizon growth and significantly delay milestones.
He emphasized the caveats. Some innovations need little compute, researchers might supply good ideas before frontier training, and AI labor could accelerate algorithmic work even without a full capability explosion. The conclusion depends on how close algorithmic progress is to being “basically determined by compute.”
METR used OpenAI’s previous tax returns and future R&D-compute projections reported by The Information, converting dollars back into FLOPs. Swyx cautioned that Meta, xAI, and DeepMind spending is largely invisible, while lab failures, consolidation, distillation, and recycled compute make the industry-level picture much less clean.
8. Prediction-market accuracy can be manufactured by participants
Asked for a naive frontier-model prior, Becker estimated 2025 time-horizon leadership at roughly 5% xAI, 50% OpenAI, and 45% Anthropic—“don’t shoot me if I’m completely incorrect”—while noting that another benchmark would produce a different distribution.
His number-one Manifold result came mostly from one charity market. Seeing donations projected linearly, he bought the higher bucket with fake currency, later made the real donation that pushed the outcome across the boundary, repeated the maneuver, then attempted a failed bluff; roughly $5,000 donated still produced enough winnings to top the leaderboard.
The hosts called it “prediction markets with high agency.” A separate best-AI-model market had attracted $28 million in trading volume even though employees might know unreleased benchmark results. Becker said Meta employees should not participate; the discussion also raised insider information as a possible source of price discovery.
Becker has become less convinced by the sector’s social-value case. Reliable probabilities on wars and other consequential events would be useful, but retail losses and “gambling-like behaviors” impose costs; a market dominated by sophisticated players trading against retail would look more worrying than large institutions trading against one another.
9. Better evaluations will watch agents fail in the open world
Becker highlighted AI Village-style tasks—launching a merchandise store, organizing a park event, or building a human-subjects experiment—as an important direction. Old and new models, vision dependence, and uncontrolled conditions complicate inference, but open-ended environments reveal how agents “fall on their face.”
His ideal test would supply broad affordances and simply instruct: “Automate R&D, go.” He expects present systems to fail through poor resource use and long-horizon execution, illustrating why excellence on a detailed software issue is “a very different thing” from automating an organization’s research process.
Agent transcripts are another enormous but selected dataset. They expose iterative actions, outputs, impressive behavior, and possible preference subversion, yet users naturally choose tasks with some chance of success; a striking unsafe transcript therefore says little by itself about the behavior’s base rate.
Benchmark success should also be separated from whether code would merge into main: did it add tests, follow repository patterns, and integrate correctly? METR tunes harnesses on development tasks and evaluates held-out tasks to limit overfitting; scaffolding has “a lot of juice,” but may be valuable within one model generation and washed away by the next.
10. METR’s 2026 roadmap pairs capability evidence with safeguards
Becker expects more time-horizon and productivity-style evidence, plus monitoring research on whether safeguards can be successfully applied to models attempting dangerous tasks. Current work is generally black-box rather than interpretability-based, feeding capability, propensity, and monitoring evidence into fuller risk assessments during 2026.
He offered no concrete 2030 success metric, but described METR as a scrappy environment where talented people work on frontier science.