Jul 24, 2026
AI

Kimi K3 cyber evaluation shows gap with leading U.S. models

UK and U.S. AI safety evaluators found Moonshot AI's Kimi K3 can assist cyber attacks but trails top U.S. models on exploits.

Colin Brandt

By Colin Brandt · Enterprise Reporter

· 3 min read

Kimi K3 cyber evaluation shows gap with leading U.S. models
Photo: The Decoder

The Kimi K3 cyber evaluation by the British AI Security Institute and the U.S. Center for AI Standards and Innovation found that Moonshot AI’s latest model can help with offensive cyber work, while still lagging well behind leading U.S. frontier systems. The result matters for AI labs and security teams because it suggests open-weight Chinese models are gaining useful attack capabilities, even if they have not matched the strongest closed U.S. models.

The institutes said Kimi K3 did not meaningfully refuse requests tied to exploit development or offensive cyber operations during the tests. They also said the model set a new mark among open-weight models and beat China’s GLM-5.2, although the gap with top U.S. systems remained wide.

How did Kimi K3 perform on cyber tests?

On ExploitBench, a Carnegie Mellon University benchmark built around 41 post-2023 vulnerabilities in Chrome’s V8 engine, Kimi K3 scored 32.2%. GLM-5.2 scored 24.4%, while the leading U.S. models averaged 76.2%, according to the institutes.

Kimi K3 failed to reach arbitrary code execution on any of the 41 ExploitBench tasks. Arbitrary code execution is the point at which an attacker can run code on a target system, making it the most serious level in this benchmark. The leading U.S. models reached that level in 20 of the 41 tasks.

The U.S. closed-weight models were tested with their system-level safeguards switched off so evaluators could measure maximum capability. Those safeguards are enabled in the public versions, according to the institutes. That distinction matters: the evaluation compares underlying capability, rather than what a normal user can obtain through public product interfaces.

What happened in the simulated network attack?

The second benchmark, called The Last Ones, tests a 32-step attack path across four subnets and about 20 hosts. The institutes said a human expert would need roughly 20 hours to finish the scenario, and that only a small set of models can complete it at all.

Kimi K3 averaged step 17 out of 32. The leading U.S. models averaged 28.5 steps, while GLM-5.2 reached 11. Kimi K3 completed the full path once in 10 attempts under a 100 million token limit, which the institutes said shows the capability exists but is not reliable. The UK AISI wrote that Kimi K3 can autonomously attack small, weakly defended and vulnerable enterprise systems when instructed and given initial network access.

The Last Ones does not include active defense, so it is not a complete stand-in for a real intrusion. For security teams, the more useful takeaway is narrower: models below the top frontier tier can already make meaningful progress through multi-step attack chains when the environment is vulnerable.

Why are distillation claims part of the Kimi K3 debate?

U.S. science adviser Michael Kratsios has accused Moonshot AI of distilling Anthropic’s Fable model by using its strongest outputs as training data for Kimi K3. He also alleged that Moonshot had access to Nvidia GB300 systems, which are subject to U.S. export controls. Moonshot’s response is not described in the evaluation.

Distillation means training one model to imitate another model’s outputs. The cyber results are consistent with one theory raised in the assessment: if Kimi K3 learned mainly from general coding, knowledge and agent outputs, it could improve on standard benchmarks without inheriting deeper offensive cyber capability that public interfaces would not expose.

CAISI’s time-series analysis shows cyber performance rising for both U.S. and Chinese models since early 2025 on an Elo-based scale, while Chinese models continue to trail U.S. systems. UK AISI previously estimated that open models were four to seven months behind frontier cyber performance, down from six to 10 months at the start of 2025. The institute warned that stronger open models create “a persistent and irreversible risk of misuse.”

This story draws on original reporting from The Decoder.

More from AI

All AI →