Chinese Models Kimi K3 and GLM-5.3 Close In on GPT-5.5 and Claude Opus 5, New Analysis Finds
A new industry analysis argues the performance gap between Chinese and Western AI models has narrowed to single-digit differences on broad benchmarks. Moonshot's Kimi K3 and Zhipu's GLM-5.3 now trail OpenAI and Anthropic's top models by only a few points on the Artificial Analysis Intelligence Index, with a clear Western edge remaining only in abstract reasoning, output reliability, and offensive cybersecurity capability.
The gap has narrowed to single digits
At launch, Moonshot AI's Kimi K3 placed third on Artificial Analysis's Intelligence Index with 57 points, trailing then-leaders GPT-5.5 and Claude Opus 4.8. Anthropic's Opus 5 subsequently retook the top spot at 61 points — a gap of just four points over K3. Alibaba's Qwen3.8-Max reportedly reaches a similar overall level.
According to the analysis, K3 improved most on agentic tasks: on AutomationBench-AA it briefly held first place until Opus 5 overtook it, and on CEO-Bench — a test where an agent runs a simulated software company for 500 days — K3 posted the best published single run at $22.15 million in simulated results. Predecessor models, including K2.7, had regularly failed these longer-horizon tasks.
One caveat: newer Chinese models reportedly consume significantly more tokens per task than Western competitors, offsetting part of their price advantage once cost-per-completed-task is factored in.
Where a Western edge still measurably exists
The analysis identifies three areas where a gap persists:
Abstract reasoning. On ARC-AGI-1, a test of abstract pattern recognition on small puzzle grids, K3 (94.5%) and Anthropic's Fable 5 (98.5%) score nearly evenly. The gap widens sharply on the harder ARC-AGI-2: 60.4% for K3 versus 89.2% for Fable 5.
Reliability. The AA-AnalystAgent benchmark, launched August 12, uses a "pass^5" metric that only counts a task solved if a model succeeds in five out of five independent runs. Opus 5 leads at 54%, GPT-5.5 at 50%, and K3 — the top open model — at 39%. Notably, K3 solves 73% of tasks at least once across five attempts, nearly matching Opus 5's 74%, indicating the gap stems from inconsistency rather than raw capability.
Cybersecurity. A joint assessment by the UK's AI Safety Institute and the US CAISI found K3 scored 32% on ExploitBench, a benchmark for developing exploits, versus roughly 76% for leading US models. K3 failed all 41 tasks requiring code execution on a target system, while US models solved 20 on average. However, GLM-5.3, released August 14, scored 54.4% on ExploitBench by Zhipu AI's own measurement — more than double its predecessor GLM-5.2 — cutting the gap to top US models roughly in half within a month. On CyberGym, a benchmark for finding and validating source-code vulnerabilities, GLM-5.3 reportedly edged past leading US models entirely.
Cybersecurity comparisons are complicated by the fact that US labs withhold their strongest capabilities from public release. Anthropic's most capable cyber model, described as Mythos 5, reaches 78% on ExploitBench but is only available under restricted access through a program called Project Glasswing; its public counterpart Fable 5 performs closer to the older Opus 4.8 level (40%) because safety filters intercepted 407 of 410 test episodes. OpenAI's GPT-5.6 Sol reportedly reaches 73.5% on ExploitBench under OpenAI's own testing but stays below the "Critical" tier on its risk framework — a threshold the company says its upcoming Astra model could cross, which would trigger stricter development controls.
What this means
The practical takeaway is that raw model-capability comparisons are becoming less useful as a business moat. If open-weight Chinese models can match closed US models within months on most broad benchmarks, the defensible value shifts away from any single model release and toward the surrounding system: proprietary data pipelines, agentic infrastructure, reliability engineering, and access-gated capabilities like advanced cyber tooling. Anthropic reportedly cites its narrowing top-tier lead to investors ahead of its IPO, but the evidence suggests that lead is retreating to a small number of specialized, harder-to-replicate areas — abstract reasoning, consistency across repeated runs, and capabilities intentionally withheld from public release — rather than disappearing outright.
Related Articles
Anthropic: Zhipu's Open-Weight GLM-5.3 Nearly Matches Claude Mythos Preview at Building Cyber Exploits
Anthropic's Frontier Red Team reports that Zhipu AI's open-weight GLM-5.3 comes close to Claude Mythos Preview on cyber exploit benchmarks, scoring 50/410 vs 56/410 on ExploitBench. Unlike Mythos Preview, GLM-5.3 shipped without effective safeguards and can be jailbroken with simple prompting tricks or abliteration.
Ramp AI Index: US business AI spending falls while usage rises about 50% from July peak
US companies are spending less on AI even as usage hit a record high at the end of September, according to the latest Ramp AI Index. Ramp economist Ara Kharazian attributes the drop almost entirely to price competition between OpenAI and Anthropic. In the last week of September, Anthropic took 51% of token spending and OpenAI 44.5%.
OpenAI says it shut down a 15,000-account campaign to extract model reasoning; the attack still worked on Azure
OpenAI says it detected and shut down an adversarial distillation campaign involving more than 15,000 accounts, linked in part to people associated with Moonshot AI. Researchers report that the same reasoning-extraction attack still worked on Microsoft Azure on September 13 against every OpenAI model they tested, including GPT-6 Astra.
OpenAI Launches Decisions API, a Fast Classifier Built on Luna Model, Echoing TypeSafe's Jev
At Dev Day, Sam Altman revealed OpenAI's new Decisions API, which narrows its Luna model to predefined choices for fast, cheap classification. The move closely mirrors Jev, a specialized decision model from startup TypeSafe AI released weeks earlier.
Comments
Loading...