
Kimi K3 Partial Mode: Prefill Replies to Force Output
Summary
Prefill Kimi K3's assistant turn to lock output shape, force clean JSON, and steer tone.
Moonshot AI shipped Kimi K3 on July 16, 2026, a 2.8-trillion-parameter open-weight model that jumped straight to the top of Hacker News and the coding-agent leaderboards. Most of the launch coverage focused on the headline numbers: 1M-token context, native vision, three reasoning-effort tiers. But the feature that quietly saves you the most headaches in production got almost no airtime: partial mode.
Partial mode lets you prefill the start of the assistant's reply and force the model to continue from that exact prefix. Instead of writing three paragraphs of prompt instructions begging the model to "respond only with JSON, no preamble, no markdown fences," you hand it the opening { and it has no choice but to keep going in that shape. It is the most reliable way to control the form of an LLM's output, and because Kimi K3's API mirrors OpenAI's Chat Completions format, it is a two-line change to your existing code.
Keep reading — it's free
Enter your email to keep reading — plus the best of AI & tech, daily. Free, forever.
Already a member? Sign in
Comments
Be the first to comment