AI NewsModels & agentsReported
Simon Willison tested Qwen3.8 27B on 5,070 addition-in-words cases and watched reasoning take it to 167 of 169
Simon Willison ran 5,070 addition-in-words cases through Qwen3.8 27B locally; the model scored 23.57 percent without reasoning and 167 of 169 with medium reasoning on.

Image: Simon Willison
Why it mattersA small open model that cannot add numbers at all can add them when reasoning is on, which changes what a cheap local model is actually good at.
A model that scores 23 percent on an arithmetic test is not supposed to score 99 percent on it an hour later.
Simon Willison published his test of Qwen3.8-27B-Q4_K_M on 4 October 2026, running the model locally on a DGX Spark and asking it to add two positive integers and give the answer in English words. He ran 5,070 cases with reasoning disabled, and a paired 169-case comparison with medium reasoning enabled. The whole experiment is on his site, with the three result heatmaps as images.
The numbers
Without reasoning, the model got the sum right 23.57 percent of the time. The number drops as the operands get longer: 97.04 percent correct for one- to three-digit addends, 6.44 percent for ten- to thirteen-digit addends. Format compliance, meaning the model answered in words at all, held at 96.17 percent across the run.
With medium reasoning turned on, the model got 167 of the 169 paired cases right. That is 98.8 percent on the harder parts of the first run.
Willison frames the test against a two-year-old GPT-4o experiment by Colin Frasier on the same task, where GPT-4o was variable but generally better than the no-reasoning Qwen number here. The point of running it today is not the comparison; it is that a small, local, cheap-to-serve open model fails arithmetic in a predictable way, and that turning on its reasoning mode fixes it in a predictable way.
The use the test actually points at
A buyer picking Qwen3.8 27B for an application has to know this difference before shipping. If the thing you need the model to do is anything that passes through numbers, which is nearly every business use (prices, counts, dates in digits, totals, scores), the no-reasoning path is going to be wrong at a rate the demo will not show you.
The fix is not free either. Reasoning adds latency and tokens; a model running inside a tight loop or an agent that cannot afford a long think step cannot pay for 169 reasoning passes on a batch of 169 small jobs. The honest answer is that the model is cheap when the task is easy and expensive when the task is hard, and this test shows exactly where the line is.
Two caveats sit on the result. 5,070 is a reasonable sample for the first number; 169 is small for the second, and Willison names it as a paired comparison rather than a repeat. And the test is one quantised GGUF on one piece of hardware; a different quantisation or a different host may move the figures around.
Source
- Qwen3.8 27B addition in words, Simon Willison, 4 October 2026.
This item was written by an AI system from the linked source. Reveneau is responsible for what it publishes.
Get AI News in your inbox
New developer tools, model and agent releases, and how teams are actually using them to release software. Short, and only when there is something worth reading.


