analysisOpenAI

Chinese Models Kimi K3 and GLM-5.3 Close In on GPT-5.5 and Claude Opus 5, New Analysis Finds

TL;DR

A new industry analysis argues the performance gap between Chinese and Western AI models has narrowed to single-digit differences on broad benchmarks. Moonshot's Kimi K3 and Zhipu's GLM-5.3 now trail OpenAI and Anthropic's top models by only a few points on the Artificial Analysis Intelligence Index, with a clear Western edge remaining only in abstract reasoning, output reliability, and offensive cybersecurity capability.

3 min read
0

The gap has narrowed to single digits

At launch, Moonshot AI's Kimi K3 placed third on Artificial Analysis's Intelligence Index with 57 points, trailing then-leaders GPT-5.5 and Claude Opus 4.8. Anthropic's Opus 5 subsequently retook the top spot at 61 points — a gap of just four points over K3. Alibaba's Qwen3.8-Max reportedly reaches a similar overall level.

According to the analysis, K3 improved most on agentic tasks: on AutomationBench-AA it briefly held first place until Opus 5 overtook it, and on CEO-Bench — a test where an agent runs a simulated software company for 500 days — K3 posted the best published single run at $22.15 million in simulated results. Predecessor models, including K2.7, had regularly failed these longer-horizon tasks.

One caveat: newer Chinese models reportedly consume significantly more tokens per task than Western competitors, offsetting part of their price advantage once cost-per-completed-task is factored in.

Where a Western edge still measurably exists

The analysis identifies three areas where a gap persists:

Abstract reasoning. On ARC-AGI-1, a test of abstract pattern recognition on small puzzle grids, K3 (94.5%) and Anthropic's Fable 5 (98.5%) score nearly evenly. The gap widens sharply on the harder ARC-AGI-2: 60.4% for K3 versus 89.2% for Fable 5.

Reliability. The AA-AnalystAgent benchmark, launched August 12, uses a "pass^5" metric that only counts a task solved if a model succeeds in five out of five independent runs. Opus 5 leads at 54%, GPT-5.5 at 50%, and K3 — the top open model — at 39%. Notably, K3 solves 73% of tasks at least once across five attempts, nearly matching Opus 5's 74%, indicating the gap stems from inconsistency rather than raw capability.

Cybersecurity. A joint assessment by the UK's AI Safety Institute and the US CAISI found K3 scored 32% on ExploitBench, a benchmark for developing exploits, versus roughly 76% for leading US models. K3 failed all 41 tasks requiring code execution on a target system, while US models solved 20 on average. However, GLM-5.3, released August 14, scored 54.4% on ExploitBench by Zhipu AI's own measurement — more than double its predecessor GLM-5.2 — cutting the gap to top US models roughly in half within a month. On CyberGym, a benchmark for finding and validating source-code vulnerabilities, GLM-5.3 reportedly edged past leading US models entirely.

Cybersecurity comparisons are complicated by the fact that US labs withhold their strongest capabilities from public release. Anthropic's most capable cyber model, described as Mythos 5, reaches 78% on ExploitBench but is only available under restricted access through a program called Project Glasswing; its public counterpart Fable 5 performs closer to the older Opus 4.8 level (40%) because safety filters intercepted 407 of 410 test episodes. OpenAI's GPT-5.6 Sol reportedly reaches 73.5% on ExploitBench under OpenAI's own testing but stays below the "Critical" tier on its risk framework — a threshold the company says its upcoming Astra model could cross, which would trigger stricter development controls.

What this means

The practical takeaway is that raw model-capability comparisons are becoming less useful as a business moat. If open-weight Chinese models can match closed US models within months on most broad benchmarks, the defensible value shifts away from any single model release and toward the surrounding system: proprietary data pipelines, agentic infrastructure, reliability engineering, and access-gated capabilities like advanced cyber tooling. Anthropic reportedly cites its narrowing top-tier lead to investors ahead of its IPO, but the evidence suggests that lead is retreating to a small number of specialized, harder-to-replicate areas — abstract reasoning, consistency across repeated runs, and capabilities intentionally withheld from public release — rather than disappearing outright.

Related Articles

product update

OpenAI Launches 'Private Safety Processing' to Detect Misuse Without Storing Enterprise Data

OpenAI has built a system called Private Safety Processing that detects misuse patterns across multiple interactions without storing customer inputs or outputs. The company says it only receives narrow safety signals—type and severity of activity—while data stays encrypted on customer infrastructure.

product update

OpenAI Previews 'Private Safety Processing' to Detect Abuse Without Retaining Customer Data

OpenAI is previewing Private Safety Processing to select customers, an automated system that monitors for misuse across multiple sessions without retaining any customer data. The move directly contrasts with Anthropic's July policy allowing 30-day data retention for 'covered models' like Fable.

analysis

Researchers Warned About Automated AI Research — Several Predicted Milestones Already Hit, New Report Says

IAPS fellow Severin Field interviewed 25 researchers from top AI labs about recursive self-improvement in late 2025. Several milestones they cited as evidence of progress — Math Olympiad gold, autonomous training loops, majority AI-written code — have since occurred, according to a new report.

product update

OpenAI Patches Codex Bug That Let AI Agent Delete Real User Files

OpenAI has shipped a security update for Codex after users reported that GPT-5.6 Sol was autonomously deleting real files instead of temporary ones. The bug stemmed from misused system variables like $HOME pointing cleanup commands at actual home directories.

Comments

Loading...

China AI Catch-Up: Kimi K3, GLM-5.3 vs GPT-5.5, Opus 5 | TPS