Pioneers Insight Method Research Author
Really Big Test-Time Compute in AI Changes Benchmarks, Safety and Research with OpenAI's Noam Brown
Back to Episodes

Really Big Test-Time Compute in AI Changes Benchmarks, Safety and Research with OpenAI's Noam Brown

Summary

  • A benchmark score without a test-time-compute budget no longer measures model quality cleanly. Brown says 5.5 looked only modestly better than 5.4 on standard grids, while users quickly found a larger improvement. He attributes the understatement to unmeasured test-time compute and describes 5.5 as more efficient with its thinking; separately, he says controlling thinking time reveals a substantial o3-over-o1 jump. The replacement is a cost, token, or time curve: “There should be an x-axis.” In practice, he says users should iterate quickly when needed and allow longer thinking when the problem warrants it.
  • The industry may be releasing models before anyone discovers their capability ceiling. Modern systems can improve for weeks and remain on an upward slope beyond 100 million tokens, while new models arrive every two or three months. Brown proposes extrapolating expensive performance from smaller runs—for example, predicting a $10,000 inference result using experiments capped at $10 or $100.
  • Safety frameworks inherit the same measurement failure, with much higher stakes. Policies designed around GPT-3 largely treat capability as intrinsic, but today it is “a function of how much money you put into it”: a model given $10,000 may do far more than at $10, and $10 million may unlock more again. Existing policies do not clearly answer which budget should govern evaluations of cyber, bioweapon, or other dangerous capabilities.
  • Brown’s poker-solver test suggests model progress is much larger than benchmark deltas imply. With 5.2 he built a river solver roughly five times faster than working alone, although the models could be unreliable; 5.5 can nearly build a full solver with gentle steering. He would not be surprised if, within six months or a year, one model could reproduce “basically my entire PhD thesis.”
  • Existing models may contain valuable scientific capability that remains uneconomic to excavate only briefly. Brown says a scaffolded 5.5 could likely have found the Erdős unit-distance disproof before OpenAI’s internal model, at a ballpark cost of $1,000–$100,000 of inference. If each release cuts such costs by 10x or 100x, and sometimes more, the investment question becomes when to deploy compute versus wait for the next cost curve.
  • Recursive self-improvement looks gradual because research taste and elapsed time remain binding constraints. Models can optimize existing algorithms by 10–100x yet still fail to invent a better one, and their strongest results require long inference runs: “Time itself becomes a bottleneck.” The larger upside may come from persistent multi-agent knowledge accumulation, not an overnight intelligence explosion.
  • Routing businesses must prove they beat simply letting the strongest model think longer at the same cost. Consensus across models can raise scores, but Brown asks whether it still wins after equalizing test-time compute—and whether benchmark gains survive real-world use. That makes budget-normalized evaluation central to assessing routing, orchestration, and model-choice layers.

Deep dive

Not yet available upstream; scheduled sync will retry.