The Hallway Track
Research Findings

Qwen3.8 27B addition in words

Simon Willison · Oct 04, 2026 · Research Findings

Qwen3.8 27B scores only 23.57% on addition-in-words tasks versus GPT-4o's far higher accuracy

“I'm confident GPT-4o didn't cheat and use a calculator, especially since it got so many of the calculations wrong”

Simon Willison replicated a GPT-4o arithmetic benchmark on local hardware using Qwen3.8-27B with reasoning disabled, finding only 23.57% overall accuracy on addition-in-words tasks — substantially worse than GPT-4o. The experiment highlights a persistent gap between frontier API models and quantized local models on structured numerical reasoning. While methodologically interesting, this is a narrow capability probe rather than a broad industry signal.

arithmetic-benchmarking Qwen3 local-models model-comparison numerical-reasoning

Watch / read the original source →