
Kimi K3 Reasoning Effort: Stream Thoughts, Cut Cost
Summary
Tune Kimi K3's low/high/max reasoning effort and stream reasoning_content to control token cost.
On July 16, 2026 Moonshot AI switched on the Kimi K3 API: a 2.8-trillion-parameter model, the first open-source system in the 3-trillion class, with full weights promised by July 27 under a modified MIT license. The headline number is the parameter count. The number that will actually show up on your invoice is reasoning tokens.
K3 is a thinking model with no off switch. Every request reasons before it answers, and those hidden reasoning tokens are billed at the output rate. That makes the one knob most people ignore — reasoning_effort — the single biggest lever on both quality and cost. Set it wrong and a batch of simple classification calls quietly burns the budget of a hard math proof.
Keep reading — it's free
Enter your email to keep reading — plus the best of AI & tech, daily. Free, forever.
Already a member? Sign in
Comments
Be the first to comment