Skip to content
NIST·

🔒Moonshot AI's Kimi K3 Falls Short in Cyber Eval

Kimi K3 Underperforms on Critical Cyber Tasks

TL;DR

UK AISI and CAISI evaluated Moonshot AI's Kimi K3, finding it underperforming compared to leading models. In cyber tasks, Kimi K3 reached only step 17 of a 32-step attack path, while others hit 28.5 steps on average.

Moonshot AI's latest model, Kimi K3, has been evaluated by UK AISI and CAISI in preliminary cyber evaluations. The results show that Kimi K3 performs significantly below the most recent frontier models when tasked with developing exploits or navigating complex attack scenarios. In 'The Last Ones' (TLO) cyber range, a 32-step simulated network attack, Kimi K3 only reached step 17 on average compared to other models hitting 28.5 steps. This matters for anyone relying on AI for cybersecurity tasks, as it highlights the limitations of current models in handling sophisticated threats.

Moonshot AI's Kimi K3 Falls Short in Cyber Eval — NIST

Key Points

1

Kimi K3 reached step 17 of the 32-step TLO cyber range in evaluations, while others hit 28.5 steps on average (July 23, 2026).

2

On ExploitBench, Kimi K3 scored 32%, outperforming GLM-5.2's 24% but failing to achieve arbitrary code execution in all tests.

3

Kimi K3's cyber capabilities were assessed using Item Response Theory (IRT) across multiple tasks and benchmarks.

4

The evaluation was conducted by UK AISI/CAISI, focusing on Kimi K3's performance against recent frontier models.

5

Despite limitations, Kimi K3 outperformed GLM-5.2 within the 100M-token limit in one of ten TLO attempts.

Why It Matters

If you're relying on AI for cybersecurity tasks, this evaluation highlights significant gaps in current model capabilities. Kimi K3's performance falls short compared to leading models, especially in complex attack scenarios and exploit development.

Moonshot AIKimi K3UK AISICAISIcyber evaluation

Frequently Asked Questions

Why does this matter?

If you're relying on AI for cybersecurity tasks, this evaluation highlights significant gaps in current model capabilities. Kimi K3's performance falls short compared to leading models, especially in complex attack scenarios and exploit development.

What happened?

UK AISI and CAISI evaluated Moonshot AI's Kimi K3, finding it underperforming compared to leading models. In cyber tasks, Kimi K3 reached only step 17 of a 32-step attack path, while others hit 28.5 steps on average.

Comments

Subscribe to join the conversation...

Be the first to comment

Enjoyed this article?

Get it daily. 7am. Free. Reads in 5 minutes.

Join 2,253 builders reading daily.

Also get