Skip to content
Qwen3.8-Flash-Next: Preserved Thinking, Cheaper Agents — ContentBuffer guide

Qwen3.8-Flash-Next: Preserved Thinking, Cheaper Agents

K
Kodetra Technologies··11 min read Intermediate

Summary

Qwen keeps reasoning across every turn by default. Exploit it, or pay for it.

Alibaba's Qwen team dropped Qwen3.8-Flash-Next on August 26, 2026 — open weights, 125B parameters with only 6B activated per token, plus a 51B n-gram embedding table and a 4B multi-token-prediction layer. It is explicitly labelled an early preview of the architecture that will underpin Qwen4, the same way Qwen3-Next previewed Qwen3.5.

The benchmark numbers got the attention: 62.5 on SWE-bench Pro, 81.0 on SWE-bench Multilingual, 73.5 on Toolathlon Verified. But the thing that will actually change your code is buried in the API section of the model card, and almost nobody is talking about it.

Keep reading — it's free

Enter your email to keep reading — plus the best of AI & tech, daily. Free, forever.

Also get
or

Already a member? Sign in

Comments

Subscribe to join the conversation...

Be the first to comment