China's Zhipu AI releases GLM-5.2, claims parity with Mythos on cybersecurity benchmarks
Zhipu AI released its open-weight GLM-5.2 model, with researchers claiming it matches Anthropic's Mythos on certain bug-finding and cybersecurity tasks. The model lags behind Anthropic and OpenAI models on general benchmarks but represents a significant narrowing of capabilities between Chinese and US AI systems.
GLM-5.2 — Quick Specs
China's Zhipu AI releases GLM-5.2, claims parity with Mythos on cybersecurity benchmarks
Zhipu AI released its open-weight GLM-5.2 model on June 28, 2026, with researchers claiming it matches Anthropic's Mythos on certain bug-finding and cybersecurity scenarios. The model represents a significant narrowing of the capability gap between Chinese and US AI systems in specialized security tasks.
Performance and capabilities
According to researchers, GLM-5.2 achieves performance comparable to Mythos on specific cybersecurity benchmarks focused on vulnerability detection and bug-finding. However, the model still lags behind Anthropic and OpenAI models on general-purpose tasks. Specific benchmark scores and pricing details have not been disclosed.
The model is released with open weights, allowing anyone to download and run it on readily available hardware without the restrictions typically imposed on frontier models.
National security implications
The development has drawn attention from US government officials who have worked to restrict China's access to advanced AI models like Anthropic's Mythos and Fable, as well as the hardware required to train and deploy them. The Trump administration has classified advanced AI models capable of identifying security vulnerabilities as national security threats.
OpenAI recently unveiled GPT-5.6, which has similarly raised concerns about potential misuse, leading to limited access controls.
Open-weight distribution concerns
Unlike restricted commercial models, GLM-5.2's open-weight release allows power users deep access and flexibility but also enables potential abuse by malicious actors who can run the model with minimal oversight. This distribution model eliminates many of the safeguards typically implemented by AI companies for sensitive capabilities.
What this means
GLM-5.2's claimed parity with Mythos on cybersecurity tasks, if verified independently, indicates China has successfully developed competitive AI capabilities in specialized domains despite US export restrictions. The open-weight release bypasses access controls that governments and companies use to prevent misuse of powerful security-focused AI systems. This creates a new dynamic where cutting-edge vulnerability detection capabilities become widely accessible, potentially benefiting both security researchers and threat actors equally.
Related Articles
Robot Safety Benchmark Finds GPT-6 Astra and Claude Fable 5.1 Rarely Refuse Dangerous Commands
A new benchmark called RoboHarm tested whether AI models controlling robotic arms would refuse dangerous commands. GPT-6 Astra completed 60 of 100 dangerous tasks and Claude Fable 5.1 completed 34, with neither model showing a reliable safety layer.
Artificial Analysis Updates Intelligence Index to v4.2, Narrows GPT-6 Astra Gap Controversy
Artificial Analysis released version 4.2 of its Intelligence Index after its original scoring showed GPT-6 Astra barely improving on its predecessor, contradicting Epoch AI's ranking of Astra as the top model out of 267 tested. The update adds two benchmarks, drops the saturated GPQA-Diamond, and increases private test weighting to 40 percent.
Zhipu AI Releases GLM-5.3, Claims It's the Strongest Open-Weights Coding Model
Zhipu AI has released GLM-5.3, a coding-focused model built on the same base as GLM-5.2 with additional post-training. The company claims it's the strongest open-weights coding model available, with gains concentrated in agentic and cybersecurity tasks, though independent benchmarks are not yet published.
Composio Benchmark: Claude Code Fastest Agent Framework, But Costs Nearly 3x More Than OpenCode
Composio benchmarked DeepSeek V4 Flash across four agent frameworks—Claude Code, Codex, OpenCode, and Oh My Pi—on 30 real-world tasks. Claude Code finished fastest at 122 seconds per task but cost $0.195, nearly three times OpenCode's $0.073, while Oh My Pi had the highest success rate at 17/30 but took 272 seconds per task.
Comments
Loading...