
Qwen3.8-Flash-Next: Preserved Thinking, Cheaper Agents
Summary
Qwen keeps reasoning across every turn by default. Exploit it, or pay for it.
Alibaba's Qwen team dropped Qwen3.8-Flash-Next on August 26, 2026 — open weights, 125B parameters with only 6B activated per token, plus a 51B n-gram embedding table and a 4B multi-token-prediction layer. It is explicitly labelled an early preview of the architecture that will underpin Qwen4, the same way Qwen3-Next previewed Qwen3.5.
The benchmark numbers got the attention: 62.5 on SWE-bench Pro, 81.0 on SWE-bench Multilingual, 73.5 on Toolathlon Verified. But the thing that will actually change your code is buried in the API section of the model card, and almost nobody is talking about it.
Keep reading — it's free
Enter your email to keep reading — plus the best of AI & tech, daily. Free, forever.
Already a member? Sign in
Comments
Be the first to comment