Mistral Large 4
Mistral Large 4 released, benchmarks criticized as saturated by frontier models
“The benchmark is saturated. Frontier models are tested with an armadillo in fishnet tights jaywalking on Mars.”
Mistral released Mistral Large 4, prompting community commentary on benchmark saturation in frontier model evaluation. Simon Willison used the release as an opportunity to run a whimsical SVG generation test across Claude, GPT, Gemini, and Mistral models. The post is more notable for the benchmark criticism meme than the model release itself.