benchmark

GLM-5.3 Ties Kimi K3 for Top Open-Model Ranking, Undercuts Rivals on Price — But Open Weights Delayed

TL;DR

Z.ai's GLM-5.3 ties Kimi K3 for the top spot among open models on the Artificial Analysis Intelligence Index, driven by a major leap in agentic task performance. The company is delaying the open-weight release by about two weeks, citing the model's unusually strong vulnerability-detection capabilities.

2 min read
0

GLM-5.3 Matches Kimi K3 at the Top of the Open-Model Field

Chinese AI startup Z.ai has released GLM-5.3, a model that ties Kimi K3 for the top spot among open models on the Artificial Analysis Intelligence Index with a score of 60. That's seven points ahead of its predecessor, GLM-5.2, marking one of the larger single-version jumps recorded for an open-weight model this year.

The biggest gains show up in agentic performance. On the GDPval-AA v2 benchmark, GLM-5.3's Elo score climbed from 1,524 to 1,770 — a 246-point increase. That places it second overall on the benchmark, trailing only Anthropic's Claude Opus 5, which scored 1,855.

Pricing Undercuts Kimi K3

Artificial Analysis estimates GLM-5.3 costs $0.68 per task, roughly 1.5 times the $0.44 cost of GLM-5.2. Despite that increase, GLM-5.3 remains 19 percent cheaper than Kimi K3, which Artificial Analysis prices at $0.84 per task. Note that this is a per-task cost estimate from Artificial Analysis's benchmark suite, not a standard per-token API rate — Z.ai has not published input/output pricing per million tokens for GLM-5.3.

According to Artificial Analysis, the combination of near-frontier intelligence scores and lower per-task cost puts GLM-5.3 in direct competition with Western frontier models on both capability and price, closing a gap that has persisted between Chinese open-weight releases and models like GPT and Claude.

Open Weights Delayed Over Security Concerns

GLM-5.3 is currently accessible only through Z.ai's API. The company says it is delaying the release of the model's open weights by roughly two weeks. According to Z.ai, the delay stems from GLM-5.3's unusually strong ability to detect security vulnerabilities — the company says it wants to strengthen internal controls and give select security partners restricted early access before making the weights broadly available.

Z.ai has not detailed what specific vulnerability-detection capabilities triggered the delay, nor has it named the security partners receiving early access. The company's stated rationale is unverified beyond its own claim.

What This Means

GLM-5.3's benchmark results suggest Chinese open-weight labs continue to close the gap with proprietary frontier models, particularly on agentic tasks where GLM-5.3 now trails only Claude Opus 5. The price advantage over Kimi K3 reinforces a pattern where open-weight competition is increasingly fought on cost-efficiency as much as raw capability.

The delayed weight release is the more unusual story. If Z.ai's stated concern about vulnerability-detection capability is accurate, it signals that some open-weight models are now capable enough at security analysis that vendors feel compelled to slow-walk public access — a reversal of the usual dynamic where speed to open release is treated as a competitive differentiator. Whether this becomes a recurring pattern for future high-capability open models, or a one-off precaution, will depend on what Z.ai's security partners find during the two-week review window.

Related Articles

benchmark

Claude Opus 5 Scores 61 on Intelligence Index, Beats Fable 5 on Cost Across Most Benchmarks

Anthropic's Claude Opus 5 posts a 61 on the Artificial Analysis Intelligence Index, narrowly beating Claude Fable 5 (60) and GPT-5.6 Sol (59) while costing less per task. The model leads in coding and knowledge-work benchmarks but shows a rising hallucination rate of 50 percent.

benchmark

Artificial Analysis Launches Optima, a Platform to Build Custom AI Benchmarks on Your Own Data

Artificial Analysis has launched Optima, a platform that lets users build custom AI benchmarks using their own data, workflows, or use-case descriptions. Unlike public benchmarks, Optima compares models on cost per task and time per task in addition to quality.

benchmark

Qwen3.8 Max Matches Claude Opus 4.8 on Intelligence Index, But Costs 2x More Per Task Than Predecessor

Alibaba's Qwen3.8 Max jumps 10 points to 56 on the Artificial Analysis Intelligence Index, putting it on par with Claude Opus 4.8. But Kimi K3 still edges it out at a lower per-task cost, and Qwen3.8 Max shows a sharp rise in hallucination rate.

benchmark

Moonshot's Kimi K3 tops Code Arena frontend benchmark at 1,679 points but scores only 39% on FrontierMath Tier 4

Moonshot AI's Kimi K3 model has claimed first place in the Code Arena frontend benchmark with a score of 1,679, surpassing Claude Fable 5 (1,631) and GPT-5.6 Sol (1,618). However, the model achieves only 39% accuracy on FrontierMath Tier 4, while top Western models from OpenAI and Anthropic reach near 90% on the same expert-level math tasks.

Comments

Loading...