HAI Seminar: Code World Models for General Game Playing
DeepMind researchers synthesize code world models that replicate games, then apply planners like MCTS instead of LLM policies.
“This code is not going to tell us how to play. It's going to tell us it's just going to replicate the game.”
A Stanford HAI seminar (with DeepMind collaborators) presents code world models, where an LLM synthesizes code that replicates a game's dynamics rather than dictating moves, allowing planners like Monte Carlo tree search or RL to run on top. It also covers hybrid LLM-plus-code systems for general game playing. This is an incremental research direction on LLM-based decision-making rather than a major industry announcement.