The Hallway Track
Governance & Policy

AI benchmark scores don’t tell you what you think they do

No Priors · Jun 29, 2026 · Governance & Policy

Current AI safety policies fail to account for test-time compute, where model capability scales with money spent.

“The capability of the model is a function of how much money you put into it, basically.”

A speaker argues that preparedness frameworks and responsible scaling policies are flawed because they evaluate a model's fixed capability, ignoring that test-time compute lets capability scale with budget. This matters because it exposes a gap in how AI safety governance defines and measures model capability in a world of inference-time scaling.

test-time-compute ai-safety benchmarks responsible-scaling evaluation

Watch / read the original source →