Why Traditional Benchmarks Fail Modern AI Models with OpenAI Research Scientist Noam Brown
Traditional single-number benchmarks fail to capture that modern AI capability scales with test-time compute budget.
“The capability of the model is a function of how much money you put into it.”
OpenAI research scientist Noam Brown argues that standard benchmark grids showing a single score per model are broken because a model's capability now depends heavily on how much test-time compute (and money) is spent at inference. This matters because existing evaluation and responsible scaling policies don't account for large-scale test-time compute, leaving open the question of at what budget models should be evaluated.