Back Home

模型評測與安全

UK–US Joint Evaluation: Kimi K3 Completes a Simulated Enterprise Intrusion, but None of 41 Exploit Attempts Achieves Arbitrary Code Execution

Preliminary testing by the UK AISI and the US CAISI shows that Kimi K3 scored 32% on ExploitBench, outperforming GLM-5.2 but trailing the tested US frontier models by a wide margin. The model also completed a 32-step simulated enterprise attack once, but the limited evaluation conditions and sample size do not support direct conclusions about real-world attack success rates.

Internet Archive Book Images · No restrictions · Image source
zh-Hant

The UK AI Security Institute and the US Center for AI Standards and Innovation have released an initial joint assessment of Kimi K3’s cybersecurity capabilities. The evaluation used 41 vulnerability exploitation tasks from ExploitBench and a simulated enterprise network called The Last Ones (TLO). Kimi K3 scored approximately 32% on ExploitBench, ahead of GLM-5.2 at 24%. However, it did not achieve the highest-level objective—arbitrary code execution—in any of the 41 samples, while the strongest model tested completed an average of 20.

ExploitBench does not simply classify outcomes as successes or failures. Instead, it breaks browser exploitation into stages including coverage, triggering a crash, creating read/write primitives, gaining control-flow control, and achieving arbitrary code execution. Kimi K3’s 32% score therefore indicates that it made progress along some attack chains; it does not mean that the model successfully compromised 32% of the targets. This also explains why it could identify promising directions but failed to reliably assemble intermediate primitives into a complete exploit.

On TLO’s 32-step attack path, Kimi K3 reached step 17 on average and completed the full sequence in one of 10 attempts within the 100-million-token limit. GLM-5.2 reached step 11 on average, while the leading model averaged step 28.5. Kimi K3’s built-in safeguards also did not prevent vulnerability development or offensive operations during testing. This is particularly noteworthy for a model expected to have open weights, because deployers cannot rely on provider-side refusals as their primary security boundary.

The report’s findings remain preliminary. Because of hosting constraints, Kimi K3 underwent only selective testing, and its overall capability estimate was based primarily on a single benchmark, ExploitBench, while other models were evaluated using a broader task set. The US closed-source models were also tested with system-level safeguards disabled. TLO has no active defender, alerting costs, or randomized network, and it intentionally preserves a viable attack path. Engineers should next watch for independent replication after the weights are released, the effects of different agent frameworks and token budgets on performance, and whether the gap between vulnerability discovery and end-to-end exploitation capabilities persists.

Sources

  1. UK AISI / CAISI Preliminary Assessment of Kimi K3's Cyber Capabilities
  2. ExploitBench: A Capability Ladder Benchmark for LLM Cybersecurity Agents
  3. Kimi K3 Tech Blog: Open Frontier Intelligence