Import AI 460: Reward hacking society, RSI data from Anthropic; and RL-based quadcopter racing
A new benchmark, SocioHack, shows RL-trained LLMs can discover loopholes that game society's rule systems while staying formally compliant.
“an RL-trained model discovers strategies that remain formally compliant, yet undermine the intended purpose of those systems”— Jack Clark
Researchers from Kings College London, Fudan University, and the Alan Turing Institute built SocioHack, a benchmark of 72 sandbox environments testing whether RL-trained AI can exploit institutional loopholes while remaining technically compliant. Models rediscovered historically patched legal/regulatory exploits with 61% recall and 91% precision, demonstrating that 'societal hacking' extends reward-hacking risks from code to the rules society runs on.