No Priors
No Priors

Really Big Test-Time Compute in AI Changes Benchmarks, Safety and Research with OpenAI Research Scientist Noam Brown

When a new AI model drops, it’s judged based on a static benchmark grid that doesn’t account for how long the model is allowed to think. How then should we measure a model’s true capability? OpenAI research scientist Noam Brown returns to talk with Sarah Guo about his latest essay on why the AI indu

Featured Speakers

Noam Brown Guest

Topics Discussed

Episode Summary

Executive Summary: Noam Brown argues that modern model evaluations are broken because they ignore test-time compute, so benchmark scores often misrepresent true capability. He says frontier models can think far longer, making performance a function of budget, time, and scaffolding. This affects model comparisons, safety evaluations, competition, and how users should deploy AI.

Main Topics: Test-time compute changes what model capability means (Priority: 5/5): Brown argues that modern model performance is no longer a fixed property; it depends heavily on how much inference budget, time, or tokens are spent during evaluation or use. Benchmark presentations are misleading (Priority: 5/5): He says standard benchmark grids hide efficiency differences and encourage shallow comparisons, because they usually omit the x-axis of compute spent, making models seem closer than they are or vice versa. Safety frameworks need a compute budget (Priority: 5/5): Responsible scaling policies and preparedness frameworks, in his view, fail to specify the amount of compute at which dangerous-capability evaluations should be run, which is now essential. Long-horizon evaluation is increasingly necessary (Priority: 4/5): Brown explains that some models can continue improving for extremely long runs—weeks or more—so some evaluations may require budgets far beyond normal release-cycle constraints. Practical use favors flexible thinking time (Priority: 4/5): He notes that while long deliberation can boost benchmark performance, real products need fast interaction and iterative workflows, so users and systems need dynamic thinking budgets. Frontier competition is accelerating research, not replacing it (Priority: 4/5): Brown believes models already speed up researchers and lab work substantially, but they are not yet producing a sudden intelligence explosion or fully automating research. Scaffolding, routing, and multi-agent systems matter—but still obey budgets (Priority: 3/5): He argues that consensus, routing, and multi-agent approaches can improve results, but they should be judged under equal compute constraints to avoid benchmark gaming.

Key Arguments: Benchmark scores without compute normalization can understate or overstate real capability; a model should be judged as a function of test-time budget. Frontier models now can keep improving over very long inference runs, sometimes weeks, so the old “run until plateau” assumption no longer works. Safety evaluations are incomplete if they do not specify or cap inference budget, because dangerous capabilities may emerge only under large compute. Some tasks, like factual recall, do not benefit much from more thinking time, while others, like Sudoku-style search, can keep improving with more compute. Model usefulness in practice often comes from fast iteration, not from making every response take a week; optimal thinking time should depend on the task. Benchmark maxing is easy through scaffolding, multi-sampling, and judge selection, so public scores can be misleading unless compute is controlled. Models are already strong collaborators for coding, optimization, and workflow support, but still lack research taste and full novel invention. Recursive self-improvement is real in a limited sense today—models accelerate researchers—but it is bottlenecked by time, compute, and non-scaling factors. Routing layers and ensemble approaches may improve outcomes, but only if compared against the same underlying compute budget as a single model thinking longer.

Data Points: GPT-3 test-time scaling: Could not meaningfully scale with $10 million of test-time compute - Used as an example of earlier models that did not benefit much from very large inference budgets. Model capability under different budgets: $10 vs. $10,000 vs. $10 million - Brown says capability is increasingly a function of inference budget, with higher budgets enabling more performance. Benchmarks with long runs: 100 million tokens - He cites AISI evaluations showing models still improving even at 100 million tokens. Potential long-horizon evaluation window: Weeks, months, or a year - Brown says some tasks may only be fully evaluated by running models for extremely long periods. Cost to explore a math problem with scaffolding: $1,000 to $100,000 - He estimates the cost of a general-purpose scaffold to solve the Erdős unit distance problem using 5.5. Model release cadence: Every 2–3 months - He says new models arrive quickly enough that full long-horizon evaluation is often overtaken by the next release. Poker solver speedup: 5x faster than alone - With 5.2, Brown says the model helped him build a river solver about five times faster. Code optimization speedup: 10x faster than he could do alone - He says the model was especially strong at optimizing code for the poker solver work. Later model improvement on optimization: 100x faster - He says later models could optimize his PhD-derived algorithms dramatically faster. Human-scale coordination: Billions of humans over 50,000 years - He contrasts AI with human civilization’s accumulated knowledge and coordination.

Pivotal Quotes: "The capability of the model is a function of how much money you put into it, basically." — Noam Brown: Core thesis of the discussion: inference budget determines realized capability. "The proper way to evaluate the models now is you either have some kind of budget for the benchmark... or you plot the performance as a function of the amount of test time compute." — Noam Brown: Explains how benchmarks should be redesigned to avoid misleading comparisons. "The problem is we're in a world now where the capability of the model is a function of how much money you put into it." — Noam Brown: Repeated emphasis on budgeted inference as a defining property of frontier systems.

Implications: Evaluations, safety checks, and product decisions should be budget-aware. The industry may need compute-normalized benchmarks, private holdouts, and dynamic inference policies to avoid misleading scores and underestimating frontier risks.

🔓 Sign Up for Unlimited Episode Search

About No Priors

View all episodes from No Priors