The Hallway Track
Industry Trends

When Will The Benchmaxxing Plague End? — Nick Heiner, Surge AI

Andrej Karpathy · AI Engineer · Aug 02, 2026 · Industry Trends

AI labs optimize for benchmark scores rather than real-world model quality, warns industry insider

“unfortunately the teams are not getting better models overall but better Elm Marina models whatever that is possibly something with a lot of nested list bullet points and emojis”— Andrej Karpathy

A Surge AI talk argues that 'benchmaxxing' — training models to score well on popular benchmarks rather than improve real-world usefulness — is endemic in the AI industry, driven by misaligned incentives and poor methodology. Andrej Karpathy is cited as observing that top-ranked LM Arena models don't match his own assessment of which models are actually best. The core problem identified is that most users lack the expertise to evaluate benchmark quality, so popularity and marketing dictate which benchmarks matter, creating a self-reinforcing feedback loop that rewards gaming over genuine improvement.

benchmarks evaluation LM Arena model quality Karpathy benchmark gaming

Watch / read the original source →