Are AI labs pelicanmaxxing?
Systematic study finds no evidence AI labs train models to draw pelicans on bicycles better
“Pelicans aren't drawn any better than other animals. Bicycles aren't drawn any better than other vehicles. And no lab draws the combination better than its pelicans and bicycles already predict.”
Dylan Castillo ran a methodical study across 7 major models testing whether AI labs specifically optimize for the 'pelican on bicycle' benchmark prompt popularized by Simon Willison. Using 48 animal-vehicle combinations run three times each, he found no statistically significant evidence of targeted benchmark gaming. The study is a fun but narrow evaluation with limited practical signal for the AI industry.