Skip to main content
●Measured savings across 11 LLMs, from Claude Opus 4.7 to Gemini Flash.→ See per-model data
Connect your client
Research

Qwen 27B fails arithmetic when forced to spell answers

A controlled experiment on Qwen3.8-27B shows the model achieves only 23.57% accuracy on addition problems when required to return answers as words rather than numerals.

1 min read

Simon Willison ran a controlled experiment on Qwen3.8-27B to measure how well the model performs addition when forced to return answers spelled out in words instead of numerals. The results reveal a stark limitation: the model achieved only 1,195 correct answers out of 5,070 test cases, for an overa...

Sign in to read the full analysis

Free account. Full analysis on LLM unit economics, plus the weekly Cost-of-Inference column.

Try it on your own context

You just read the writeup. Now run the thing. Paste a doc or some verbose tool output and watch it shrink — free, no signup.

2,912/12,000 chars
Compressed
Compressed text will appear here…
Method & sources
Source type
Primary publication (lab/vendor blog) — our analysis + implication
Source link
Simon Willison
Published
UTC
Byline
By the gotcontext.ai team (editorial standards)
Correction?
corrections@gotcontext.ai

Related