benchmarkMoonshot AI

Kimi K3 Scores 32% on Cyber Exploit Benchmark vs. 76% for Leading U.S. Models, Joint UK-US Study Finds

TL;DR

A joint evaluation by the UK AI Security Institute and U.S. Center for AI Standards and Innovation found Kimi K3 scores 32.2% on the ExploitBench benchmark versus 76.2% for leading U.S. models, though it beats China's GLM-5.2 at 24.4%. The gap may stem from Moonshot AI distilling Claude outputs that exclude advanced offensive cyber content.

3 min read
0

Kimi K3 Falls Short on Offensive Cyber Benchmarks

Moonshot AI's Kimi K3 scored 32.2% on ExploitBench, a Carnegie Mellon-developed benchmark measuring exploit development skills, compared to 76.2% for leading U.S. frontier models, according to a joint evaluation by the British AI Security Institute (UK AISI) and the U.S. Center for AI Standards and Innovation (CAISI). Kimi K3 did outperform China's GLM-5.2, which scored 24.4%, setting a new high mark among open-weight models.

The evaluation found that Kimi K3's safeguards did not meaningfully block exploit development or offensive cyber operations — the model assisted with both without resistance.

Exploit and Network Attack Results

ExploitBench uses 41 vulnerabilities discovered in Chrome's V8 engine after 2023 to measure how far a model can progress through the software exploitation process. Kimi K3 failed to reach the highest severity level — Arbitrary Code Execution (ACE), which grants full control of a target system — on any of the 41 tasks. Leading U.S. models achieved ACE on 20 of 41 tasks. The U.S. models were tested with system-level safeguards disabled to measure maximum capability; those safeguards remain active in publicly available versions.

A second test, "The Last Ones" (TLO), simulates a 32-step corporate network attack across four subnets and roughly 20 hosts — a task requiring an estimated 20 hours for a human expert, according to the institutes. Kimi K3 averaged 17 of 32 steps, versus 28.5 for leading U.S. models and 11 for GLM-5.2. Kimi K3 completed the full attack path in one of ten attempts within a 100-million-token limit, indicating the capability exists but isn't reliably accessible. AISI noted Kimi K3 "is capable of autonomously attacking small, weakly defended and vulnerable enterprise systems, when directed to do so and given initial network access."

The China-U.S. Cyber Capability Gap

CAISI's Elo-based time-series analysis, tracking model cyber capability since early 2025, shows both U.S. and Chinese models improving, with Chinese systems consistently trailing. UK AISI previously estimated the gap for open models at four to seven months, down from six to ten months at the start of 2025. The institute cautioned against complacency, warning that growing cyber capabilities in open-weight models create "a persistent and irreversible risk of misuse."

Distillation Allegations Add Context

The results align with recent distillation allegations against Moonshot AI. U.S. science advisor Michael Kratsios has accused the company of distilling Anthropic's Claude models — using Claude's outputs as training data — and separately alleged Moonshot AI accessed export-controlled Nvidia GB300 chips.

One explanation offered for Kimi K3's benchmark pattern: if the model was trained heavily on Claude outputs for general knowledge, coding, and agent tasks, it would inherit little advanced cyber capability, since Anthropic's safety classifiers block advanced offensive cyber queries in Claude's public outputs. That would let Kimi K3 match Western models on standard benchmarks while lagging significantly on cyber tasks — precisely the pattern AISI observed, since the U.S. models' cyber capabilities were only exposed after researchers disabled safeguards unavailable through public interfaces.

What This Means

This evaluation is the clearest evidence yet that raw benchmark parity between Chinese and U.S. open models doesn't imply equal capability across all domains — cyber offense in particular remains a laggard for distilled models. If Kimi K3's gap traces back to training on safety-filtered Claude outputs, it suggests distillation as a strategy has a ceiling: models can approximate a teacher's general competence but not capabilities the teacher deliberately withholds. That's a meaningful constraint on how much smaller labs can catch up to frontier safety-conscious developers purely through distillation, though the trend lines show the gap narrowing over time regardless.

Related Articles

model release

Moonshot AI Open-Sources Kimi K3 Weights After Model Matched GPT-5.6 Sol on Benchmarks

Moonshot AI has released open weights, a technical report, and supporting infrastructure for Kimi K3, a model that claims 2.5x more intelligence per unit of compute. Independent testing found notable gaps in cybersecurity and math performance compared to Western frontier models.

benchmark

Qwen3.8 Max Matches Claude Opus 4.8 on Intelligence Index, But Costs 2x More Per Task Than Predecessor

Alibaba's Qwen3.8 Max jumps 10 points to 56 on the Artificial Analysis Intelligence Index, putting it on par with Claude Opus 4.8. But Kimi K3 still edges it out at a lower per-task cost, and Qwen3.8 Max shows a sharp rise in hallucination rate.

model release

Moonshot AI Releases Kimi K3, a 2.8 Trillion Parameter Open-Weight Model; AWS Publishes Deployment Guide

Moonshot AI released Kimi K3 on July 27, 2026, a 2.8 trillion parameter Mixture-of-Experts model with a 1 million token context window and native multimodal support. AWS has published a deployment guide covering SageMaker HyperPod and Amazon EKS using ml.p6-b300.48xlarge instances with 8 NVIDIA B300 Blackwell Ultra GPUs.

model release

Moonshot AI Releases Kimi K3 Weights: 2.8 Trillion Parameters, Tighter Commercial License

Moonshot AI has released the weights for Kimi K3, a 2.8 trillion parameter model weighing in at 1.56TB on Hugging Face. The new license drops the 'modified MIT' framing and now requires companies earning over $20 million in 12-month revenue from Model-as-a-Service offerings to sign a separate agreement with Moonshot.

Comments

Loading...