The Best Models Still Reason Like Toddlers — Andrew Dai, Elorian
Today's top vision models fail at basic visual reasoning, relying on pattern recognition over spatial understanding.
“there's actually a big gap between how these models handle visual thinking and how humans do it”
The speaker demonstrates that frontier models (Claude, ChatGPT, Gemini) systematically fail at visual reasoning tasks like counting chessboard squares, counting Catan roads, and tracking robot manipulation, because they lean on pattern recognition rather than genuine spatial understanding. This matters because it exposes a large gap between current multimodal capabilities and human visual thinking, with direct implications for robotics and any application requiring reliable spatial perception.