Training Frontier Models to Out-Think Hackers — Uri Rolls, Arithmetic & Thom Wolf, Hugging Face
New cybersecurity benchmark exposes AI world-modeling gap with 1-2% model success rates
“models have one to two% success rate on this generic benchmark and the reason is the current model even though they're really good they can't really build a dynamic model of what's happening in the world”
Arithmetic and Hugging Face's Thom Wolf presented a new cybersecurity-focused AI benchmark that tests dynamic world-model construction, claiming it rivals ARC-AGI 3 in difficulty with frontier models achieving only 1-2% success rates. The benchmark requires models to infer causal state changes in interactive environments, exposing a fundamental limitation in current AI reasoning. The speakers also challenged the binary closed/open-source narrative in AI security, arguing open-source models are part of the cybersecurity solution rather than a liability.