benchmarkZhipu AI

China's Zhipu AI releases GLM-5.2, claims parity with Mythos on cybersecurity benchmarks

TL;DR

Zhipu AI released its open-weight GLM-5.2 model, with researchers claiming it matches Anthropic's Mythos on certain bug-finding and cybersecurity tasks. The model lags behind Anthropic and OpenAI models on general benchmarks but represents a significant narrowing of capabilities between Chinese and US AI systems.

2 min read
0

GLM-5.2 — Quick Specs

Context window1000K tokens
Input$0.826/1M tokens
Output$2.596/1M tokens

China's Zhipu AI releases GLM-5.2, claims parity with Mythos on cybersecurity benchmarks

Zhipu AI released its open-weight GLM-5.2 model on June 28, 2026, with researchers claiming it matches Anthropic's Mythos on certain bug-finding and cybersecurity scenarios. The model represents a significant narrowing of the capability gap between Chinese and US AI systems in specialized security tasks.

Performance and capabilities

According to researchers, GLM-5.2 achieves performance comparable to Mythos on specific cybersecurity benchmarks focused on vulnerability detection and bug-finding. However, the model still lags behind Anthropic and OpenAI models on general-purpose tasks. Specific benchmark scores and pricing details have not been disclosed.

The model is released with open weights, allowing anyone to download and run it on readily available hardware without the restrictions typically imposed on frontier models.

National security implications

The development has drawn attention from US government officials who have worked to restrict China's access to advanced AI models like Anthropic's Mythos and Fable, as well as the hardware required to train and deploy them. The Trump administration has classified advanced AI models capable of identifying security vulnerabilities as national security threats.

OpenAI recently unveiled GPT-5.6, which has similarly raised concerns about potential misuse, leading to limited access controls.

Open-weight distribution concerns

Unlike restricted commercial models, GLM-5.2's open-weight release allows power users deep access and flexibility but also enables potential abuse by malicious actors who can run the model with minimal oversight. This distribution model eliminates many of the safeguards typically implemented by AI companies for sensitive capabilities.

What this means

GLM-5.2's claimed parity with Mythos on cybersecurity tasks, if verified independently, indicates China has successfully developed competitive AI capabilities in specialized domains despite US export restrictions. The open-weight release bypasses access controls that governments and companies use to prevent misuse of powerful security-focused AI systems. This creates a new dynamic where cutting-edge vulnerability detection capabilities become widely accessible, potentially benefiting both security researchers and threat actors equally.

Related Articles

benchmark

Robot Safety Benchmark Finds GPT-6 Astra and Claude Fable 5.1 Rarely Refuse Dangerous Commands

A new benchmark called RoboHarm tested whether AI models controlling robotic arms would refuse dangerous commands. GPT-6 Astra completed 60 of 100 dangerous tasks and Claude Fable 5.1 completed 34, with neither model showing a reliable safety layer.

benchmark

Artificial Analysis Updates Intelligence Index to v4.2, Narrows GPT-6 Astra Gap Controversy

Artificial Analysis released version 4.2 of its Intelligence Index after its original scoring showed GPT-6 Astra barely improving on its predecessor, contradicting Epoch AI's ranking of Astra as the top model out of 267 tested. The update adds two benchmarks, drops the saturated GPQA-Diamond, and increases private test weighting to 40 percent.

model release

Zhipu AI Releases GLM-5.3, Claims It's the Strongest Open-Weights Coding Model

Zhipu AI has released GLM-5.3, a coding-focused model built on the same base as GLM-5.2 with additional post-training. The company claims it's the strongest open-weights coding model available, with gains concentrated in agentic and cybersecurity tasks, though independent benchmarks are not yet published.

benchmark

Composio Benchmark: Claude Code Fastest Agent Framework, But Costs Nearly 3x More Than OpenCode

Composio benchmarked DeepSeek V4 Flash across four agent frameworks—Claude Code, Codex, OpenCode, and Oh My Pi—on 30 real-world tasks. Claude Code finished fastest at 122 seconds per task but cost $0.195, nearly three times OpenCode's $0.073, while Oh My Pi had the highest success rate at 17/30 but took 272 seconds per task.

Comments

Loading...