The Hallway Track
Research Findings

Why Traditional Benchmarks Fail Modern AI Models with OpenAI Research Scientist Noam Brown

No Priors · Jun 26, 2026 · Research Findings

Traditional single-number benchmarks fail to capture that modern AI capability scales with test-time compute budget.

“The capability of the model is a function of how much money you put into it.”

OpenAI research scientist Noam Brown argues that standard benchmark grids showing a single score per model are broken because a model's capability now depends heavily on how much test-time compute (and money) is spent at inference. This matters because existing evaluation and responsible scaling policies don't account for large-scale test-time compute, leaving open the question of at what budget models should be evaluated.

test-time-compute benchmarks evaluation inference-scaling reasoning-models

Watch / read the original source →