Kimi K3 Scores 32% on Cyber Exploit Benchmark vs. 76% for Leading U.S. Models, Joint UK-US Study Finds
A joint evaluation by the UK AI Security Institute and U.S. Center for AI Standards and Innovation found Kimi K3 scores 32.2% on the ExploitBench benchmark versus 76.2% for leading U.S. models, though it beats China's GLM-5.2 at 24.4%. The gap may stem from Moonshot AI distilling Claude outputs that exclude advanced offensive cyber content.
Kimi K3 Falls Short on Offensive Cyber Benchmarks
Moonshot AI's Kimi K3 scored 32.2% on ExploitBench, a Carnegie Mellon-developed benchmark measuring exploit development skills, compared to 76.2% for leading U.S. frontier models, according to a joint evaluation by the British AI Security Institute (UK AISI) and the U.S. Center for AI Standards and Innovation (CAISI). Kimi K3 did outperform China's GLM-5.2, which scored 24.4%, setting a new high mark among open-weight models.
The evaluation found that Kimi K3's safeguards did not meaningfully block exploit development or offensive cyber operations — the model assisted with both without resistance.
Exploit and Network Attack Results
ExploitBench uses 41 vulnerabilities discovered in Chrome's V8 engine after 2023 to measure how far a model can progress through the software exploitation process. Kimi K3 failed to reach the highest severity level — Arbitrary Code Execution (ACE), which grants full control of a target system — on any of the 41 tasks. Leading U.S. models achieved ACE on 20 of 41 tasks. The U.S. models were tested with system-level safeguards disabled to measure maximum capability; those safeguards remain active in publicly available versions.
A second test, "The Last Ones" (TLO), simulates a 32-step corporate network attack across four subnets and roughly 20 hosts — a task requiring an estimated 20 hours for a human expert, according to the institutes. Kimi K3 averaged 17 of 32 steps, versus 28.5 for leading U.S. models and 11 for GLM-5.2. Kimi K3 completed the full attack path in one of ten attempts within a 100-million-token limit, indicating the capability exists but isn't reliably accessible. AISI noted Kimi K3 "is capable of autonomously attacking small, weakly defended and vulnerable enterprise systems, when directed to do so and given initial network access."
The China-U.S. Cyber Capability Gap
CAISI's Elo-based time-series analysis, tracking model cyber capability since early 2025, shows both U.S. and Chinese models improving, with Chinese systems consistently trailing. UK AISI previously estimated the gap for open models at four to seven months, down from six to ten months at the start of 2025. The institute cautioned against complacency, warning that growing cyber capabilities in open-weight models create "a persistent and irreversible risk of misuse."
Distillation Allegations Add Context
The results align with recent distillation allegations against Moonshot AI. U.S. science advisor Michael Kratsios has accused the company of distilling Anthropic's Claude models — using Claude's outputs as training data — and separately alleged Moonshot AI accessed export-controlled Nvidia GB300 chips.
One explanation offered for Kimi K3's benchmark pattern: if the model was trained heavily on Claude outputs for general knowledge, coding, and agent tasks, it would inherit little advanced cyber capability, since Anthropic's safety classifiers block advanced offensive cyber queries in Claude's public outputs. That would let Kimi K3 match Western models on standard benchmarks while lagging significantly on cyber tasks — precisely the pattern AISI observed, since the U.S. models' cyber capabilities were only exposed after researchers disabled safeguards unavailable through public interfaces.
What This Means
This evaluation is the clearest evidence yet that raw benchmark parity between Chinese and U.S. open models doesn't imply equal capability across all domains — cyber offense in particular remains a laggard for distilled models. If Kimi K3's gap traces back to training on safety-filtered Claude outputs, it suggests distillation as a strategy has a ceiling: models can approximate a teacher's general competence but not capabilities the teacher deliberately withholds. That's a meaningful constraint on how much smaller labs can catch up to frontier safety-conscious developers purely through distillation, though the trend lines show the gap narrowing over time regardless.
Related Articles
Moonshot's Kimi K3 tops Code Arena frontend benchmark at 1,679 points but scores only 39% on FrontierMath Tier 4
Moonshot AI's Kimi K3 model has claimed first place in the Code Arena frontend benchmark with a score of 1,679, surpassing Claude Fable 5 (1,631) and GPT-5.6 Sol (1,618). However, the model achieves only 39% accuracy on FrontierMath Tier 4, while top Western models from OpenAI and Anthropic reach near 90% on the same expert-level math tasks.
Moonshot AI's Kimi K3 ranks #2 globally, will release 2.8T parameter weights July 27
Moonshot AI released Kimi K3 on July 16, 2026, a 2.8 trillion parameter mixture-of-experts model that ranks #2 on the Vals AI index and #3 on Artificial Analysis's Intelligence Index. The company will release the model's weights on July 27, making it the strongest open-weight model to date, surpassing all previous open releases including DeepSeek R1.
Moonshot AI and Alibaba release 2.8T and 2.4T parameter models, claim performance near GPT-5.6 and Claude Fable 5
Within days, Moonshot AI and Alibaba unveiled what they claim are frontier-class models. Moonshot's Kimi K3, at 2.8 trillion parameters, and Alibaba's Qwen3.8, at 2.4 trillion parameters, will both be released as open-weight models with full weights available for download.
Moonshot AI releases Kimi K3, largest open-weight model at 2.8 trillion parameters
Moonshot AI released Kimi K3 on July 16, 2025, an open-weight model with 2.8 trillion parameters. The model represents the largest openly available model by parameter count, entering what the industry categorizes as the 3T class.
Comments
Loading...