AI benchmark scores don’t tell you what you think they do
Current AI safety policies fail to account for test-time compute, where model capability scales with money spent.
“The capability of the model is a function of how much money you put into it, basically.”
A speaker argues that preparedness frameworks and responsible scaling policies are flawed because they evaluate a model's fixed capability, ignoring that test-time compute lets capability scale with budget. This matters because it exposes a gap in how AI safety governance defines and measures model capability in a world of inference-time scaling.