The number that showed up in Simon Willison's feed on 17 Aug 2026 is 52. Artificial Analysis Intelligence Index. Same score they gave GPT-5.6 Luna (max). One point behind GLM-5.2 (max) and DeepSeek V4 Pro 0813 (max). That GLM is 753B. The DeepSeek is 1.7T. Qwen 3.8 27B is 27 billion parameters, Apache 2, vision in, about 17GB as a Q4_K_M GGUF.
I sat with that. A comparable open-weight bucket median of 9, and this dense 27B is sitting next to frontier max settings. I still do not treat an index as an acceptance test. I do treat it as a reason to stop saying local is for toys.

Redrawn from the Artificial Analysis Qwen3.8 27B page using the scores Willison listed on 17 Aug 2026. Higher is better. AA also says the eval burned 160M output tokens against a 43M median.
What we change on day one
Willison's 16 Aug write-up is the one I send before anyone clicks Download. Alibaba's docs default reasoning_effort to xhigh. The LM Studio GGUF kept that default. He asked for a pelican on a bike and waited 21 minutes while the model spent 22,276 reasoning tokens. A circle prompt turned into a geometric study. The weights are fine. The default will melt an 8k context window on a Tuesday.
We pin low or off. medium stays for the one job that actually needs a trace. We log token counts. Custom software development services is that pin, the tool-call path, and an off switch.
The sibling post is thirty billion on the desk, three billion awake. Nemotron is sparse. This Qwen is dense. Different VRAM math. Same product question: what may it read, what may it call, where the transcript goes.
What we still will not promise
That 52 stays 52 next month. Indexes move. That 15 to 30 tokens a second replaces hosted Luna for a chat widget. Willison saw that speed on an M5 Max and a DGX Spark. Fine for overnight jobs.
llama.cpp already has MTP draft speculation for this GGUF. I have not timed it on our boxes. I will not quote a speedup I have not measured.
If the board question is residency versus API, that is open weights are a procurement question. This post is the 17GB file and the default that eats the context.
We install the GGUF, pin the effort, and attach it to the product you already run. Request a quote if the current plan is drop Qwen into LM Studio and call it a copilot.
I still want our own eval set on your tickets before I call 52 a ship number. The index is public. Your workflow is not.
Sources
- Qwen 3.8 27B scores 52 on the Artificial Analysis Intelligence Index — Simon Willison, 17 Aug 2026
- Qwen 3.8 27B is excellent, but it defaults to wildly overthinking things — Simon Willison, 16 Aug 2026
- Qwen3.8 27B Intelligence, Performance & Price Analysis — Artificial Analysis, fetched 19 Aug 2026

