Qwen3.8 27B addition in words
Qwen3.8 27B scores only 23.57% on addition-in-words tasks versus GPT-4o's far higher accuracy
“I'm confident GPT-4o didn't cheat and use a calculator, especially since it got so many of the calculations wrong”
Simon Willison replicated a GPT-4o arithmetic benchmark on local hardware using Qwen3.8-27B with reasoning disabled, finding only 23.57% overall accuracy on addition-in-words tasks — substantially worse than GPT-4o. The experiment highlights a persistent gap between frontier API models and quantized local models on structured numerical reasoning. While methodologically interesting, this is a narrow capability probe rather than a broad industry signal.