Research
Qwen 27B fails arithmetic when forced to spell answers
A controlled experiment on Qwen3.8-27B shows the model achieves only 23.57% accuracy on addition problems when required to return answers as words rather than numerals.
1 min read
SourceSimon Willison
Simon Willison ran a controlled experiment on Qwen3.8-27B to measure how well the model performs addition when forced to return answers spelled out in words instead of numerals. The results reveal a stark limitation: the model achieved only 1,195 correct answers out of 5,070 test cases, for an overa...
Sign in to read the full analysis
Free account. Full analysis on LLM unit economics, plus the weekly Cost-of-Inference column.
Try it on your own context
You just read the writeup. Now run the thing. Paste a doc or some verbose tool output and watch it shrink — free, no signup.
2,912/12,000 chars
Compressed
Compressed text will appear here…
Method & sources
- Source type
- Primary publication (lab/vendor blog) — our analysis + implication
- Source link
- Simon Willison
- Published
- UTC
- Byline
- By the gotcontext.ai team (editorial standards)
- Correction?
- corrections@gotcontext.ai