benchmark

Zhipu's GLM-5.2 matches Anthropic's Claude Opus 4.8 on agentic benchmark at one-fifth the cost

TL;DR

Zhipu AI's open-source GLM-5.2 model scores within one percentage point of Anthropic's Claude Opus 4.8 on a key agentic benchmark while costing approximately one-fifth as much. The release comes as U.S. government restrictions limit access to Anthropic's Fable and OpenAI's GPT-5.6 models.

2 min read
0

GLM-5.2 — Quick Specs

Context window1000K tokens
Input$0.826/1M tokens
Output$2.596/1M tokens

Chinese open-source model challenges frontier labs on price-performance

Zhipu AI's GLM-5.2 scores within one percentage point of Anthropic's Claude Opus 4.8 on agentic benchmarks while costing roughly 20% as much, according to the Chinese AI startup. The open-source model, released last week, has surpassed all other open releases on the closely watched benchmark.

OpenRouter token traffic for GLM-5.2 is climbing faster than it did following DeepSeek's V4 launch in April, suggesting rapid developer adoption. Unlike DeepSeek, which focused primarily on chat applications, GLM-5.2 demonstrates strength in agentic tasks including planning, coding, testing, and task looping—capabilities enterprises are prioritizing for automation.

Enterprise cost pressures drive adoption

"I've been consistently surprised by how quickly the open source has caught up," Gabe Pereyra, co-founder of legal AI company Harvey, told CNBC. "GLM-5.2, you're seeing the first model where it's really competitive with some of these closed-source frontier models."

As token spend strains enterprise AI budgets, "intelligence per dollar" is emerging as the critical metric. GLM-5.2's combination of competitive performance and significantly lower costs addresses this pressure directly. The model is free to download, fine-tune, and deploy on enterprise infrastructure, eliminating recurring API costs.

U.S. restrictions create opening for open source

The timing coincides with increased U.S. government oversight of frontier models. Anthropic pulled its Fable Mythos-class model following a Trump administration order. OpenAI announced Friday it is limiting access to GPT-5.6 models "at the request of the U.S. government."

These restrictions make models that "no one can revoke" more attractive to enterprises concerned about deployment stability, according to the report. Open-source models eliminate dependency on external API access and regulatory compliance from third-party providers.

Specific benchmark numbers not disclosed

While the article states GLM-5.2 lands "within a percentage point" of Claude Opus 4.8 on "a key agentic benchmark," neither the specific benchmark name nor exact scores are disclosed. The pricing comparison—"roughly a fifth of the cost"—also lacks precise per-token figures for verification.

Zhipu AI, based in Beijing, has been building foundation models since 2019. The company previously released GLM-4 and other models in its series.

What this means

If verified, GLM-5.2's performance represents a significant compression in the gap between open-source and frontier closed models, particularly for agentic workflows. The combination of regulatory uncertainty around U.S. models and budget pressures on token spend could accelerate enterprise adoption of open-source alternatives. However, the lack of disclosed benchmark specifics makes it difficult to independently verify the claimed parity with Claude Opus 4.8. The shift toward "intelligence per dollar" as the primary evaluation metric reflects a maturing market moving beyond pure capability races.

Source: cnbc.com

Related Articles

benchmark

Artificial Analysis Updates Intelligence Index to v4.2, Narrows GPT-6 Astra Gap Controversy

Artificial Analysis released version 4.2 of its Intelligence Index after its original scoring showed GPT-6 Astra barely improving on its predecessor, contradicting Epoch AI's ranking of Astra as the top model out of 267 tested. The update adds two benchmarks, drops the saturated GPQA-Diamond, and increases private test weighting to 40 percent.

benchmark

Claude Opus 5 Scores 61 on Intelligence Index, Beats Fable 5 on Cost Across Most Benchmarks

Anthropic's Claude Opus 5 posts a 61 on the Artificial Analysis Intelligence Index, narrowly beating Claude Fable 5 (60) and GPT-5.6 Sol (59) while costing less per task. The model leads in coding and knowledge-work benchmarks but shows a rising hallucination rate of 50 percent.

benchmark

GLM-5.3 Ties Kimi K3 for Top Open-Model Ranking, Undercuts Rivals on Price — But Open Weights Delayed

Z.ai's GLM-5.3 ties Kimi K3 for the top spot among open models on the Artificial Analysis Intelligence Index, driven by a major leap in agentic task performance. The company is delaying the open-weight release by about two weeks, citing the model's unusually strong vulnerability-detection capabilities.

benchmark

Artificial Analysis Launches Optima, a Platform to Build Custom AI Benchmarks on Your Own Data

Artificial Analysis has launched Optima, a platform that lets users build custom AI benchmarks using their own data, workflows, or use-case descriptions. Unlike public benchmarks, Optima compares models on cost per task and time per task in addition to quality.

Comments

Loading...