The Hallway Track
Research Findings

Import AI 469: Science AI; RSI simulator; and Zuck's technological pessimism

Mark Zuckerberg · Import AI · Aug 17, 2026 · Research Findings

DiG-bench shows Fable 5 displays creative intuition; current frontier models cannot beat discovery games

“The new frontier for analyzing AI systems is understanding how good they are at inferring the unwritten rules of their environment”— Mark Zuckerberg

DiG-bench is a new 70-game benchmark from an Oxford/MIT/Princeton consortium testing AI systems' ability to infer hidden rules through exploration rather than instruction, with all games remaining unbeatable by today's frontier models. Fable 5 with Claude Code is singled out as showing early creative intuition on the benchmark, a notable callout for Anthropic's newest model. The benchmark uses private, handcrafted games to prevent training contamination, making it a more robust evaluation signal than typical public leaderboards.

benchmark AI evaluation Fable discovery creative reasoning game-playing

Watch / read the original source →