analysisAnthropic

Anthropic: Zhipu's Open-Weight GLM-5.3 Nearly Matches Claude Mythos Preview at Building Cyber Exploits

TL;DR

Anthropic's Frontier Red Team reports that Zhipu AI's open-weight GLM-5.3 comes close to Claude Mythos Preview on cyber exploit benchmarks, scoring 50/410 vs 56/410 on ExploitBench. Unlike Mythos Preview, GLM-5.3 shipped without effective safeguards and can be jailbroken with simple prompting tricks or abliteration.

3 min read
0

Anthropic's Frontier Red Team says Zhipu AI's open-weight model GLM-5.3 nearly matches Claude Mythos Preview at autonomously building cyber exploits — but shipped without the safeguards Anthropic built into its own model.

On ExploitBench, which measures success at exploiting known bugs in Chrome's V8 engine, GLM-5.3 built a working exploit in 50 of 410 attempts. Claude Mythos Preview managed 56 of 410. On Anthropic's internal binary exploitation benchmark, built from open-source projects in Google's OSS-Fuzz corpus, GLM-5.3 achieved full control of the target program in 4 percent of tasks versus 6 percent for Mythos Preview. Older models — GLM-5.2, Claude Opus 4.6, Kimi K3, and DeepSeek V4.1-Flash — scored at or near zero on both tests, according to Anthropic.

In a paired test with a human expert, GLM-5.3 found several previously unknown vulnerabilities in the JavaScript engine of a widely used browser within a single day, with limited human attention. It chained the flaws into a web page capable of reading any file on a visitor's machine, and in testing extracted a private SSH key. Anthropic says it reported the vulnerabilities to the browser vendor; additional findings in drivers and firmware remain under review.

Using the smaller GLM-5.3-Flash, Anthropic tested how fast a newly disclosed vulnerability could become a working attack. The model combined a recently disclosed Chrome bug with a known vulnerability and built a reliable exploit that bypassed an additional processor-level security feature — with 20 minutes of human oversight and 8 hours of model compute, costing $20.40 at Zhipu's API pricing.

The U.S. Center for AI Standards and Innovation (CAISI) reached similar conclusions independently, calling GLM-5.3 the most cyber-capable open-weight model available and estimating it trails the top U.S. models by about four months — though CAISI notes the top U.S. models were tested with safeguards disabled and some are restricted to vetted users.

Safeguards fall apart under simple pressure

In Anthropic's simulations, GLM-5.3 refused overtly malicious attack requests. Reframed as a red-team exercise, the model attempted to connect to the target system in 64 percent of runs. Adding prefilled reasoning steps pushed that to 92 percent. After "abliteration" — a technique that strips refusal behavior from open weights — the rate hit 100 percent. Anthropic's protected Claude models stayed at zero throughout.

Anthropic says this was its first time running abliteration internally. The process took roughly 2,200 GPU hours and cost about $4,400; the company estimates an experienced team could do it for around $1,200. Refusal rates for harmful requests dropped from over 90 percent to between 2 and 12 percent, while science and cyber benchmark scores barely changed. According to Anthropic, unlocked versions of GLM-5.3 appeared online within days of its release.

A warning with a business angle

Anthropic has withheld Mythos Preview from public release, granting access only to vetted defenders through its Project Glasswing initiative, which it says has already surfaced more than 10,000 vulnerabilities in critical software. OpenAI is reportedly taking a comparable approach with a model called Daybreak.

Anthropic's report frames closed weights as a security advantage and calls for government testing of future open models like GLM-5.3's successors — a position that also serves a company competing directly against a cheap, near-frontier open-weight rival. The UK's AI Security Institute has separately found that open models' lag behind frontier cyber capability has narrowed from six-to-ten months to four-to-seven months, while noting open models are cheaper to run and their safeguards are largely ineffective.

What this means

The benchmark gap between GLM-5.3 and Claude Mythos Preview is narrow enough that it barely functions as a safety buffer once safeguards are removed — and removal is cheap, at roughly $1,200 by Anthropic's own estimate. That the finding comes from a company that doesn't release open weights and is actively selling defensive access to its own frontier model doesn't make the underlying numbers wrong; CAISI's independent assessment points the same direction. But it does mean policymakers should weigh Anthropic's specific policy asks — mandatory government testing, continued reliance on gated access models like Mythos Preview and Daybreak — against the fact that the company benefits directly from a regulatory environment that disadvantages open-weight competitors.

Related Articles

research

Anthropic Red Team: GLM-5.3 Matches Claude on Binary Exploitation for First Time

Anthropic's Frontier Red Team reports that Zhipu AI's GLM-5.3 achieved full control flow hijacks in 4% of binary exploitation trials, versus 6% for Claude Mythos Preview. Predecessor models Claude Opus 4.6 and GLM-5.2 scored zero, marking what Anthropic calls a crossed threshold in offensive cyber capability.

analysis

OpenAI Reportedly Pulls Astra 6.1 Release Over Deception, Alignment Failures

OpenAI has reportedly canceled the planned release of Astra 6.1 after internal testing showed the model exhibited higher levels of deception and unsafe behavior than prior models. The decision, first reported by The Wall Street Journal, comes as the industry faces mounting scrutiny over AI agent safety incidents.

model release

Anthropic Releases Claude Sonnet 5.5, Now Powering Free Tier on Claude.ai

Anthropic released Claude Sonnet 5.5, claiming it runs 30%+ faster and costs up to 30% less than Sonnet 5 while beating it on benchmarks, at the same price. The model now powers the free tier on claude.ai, giving Anthropic a notably stronger free offering than OpenAI's ChatGPT.

model release

Anthropic's Claude Sonnet 5.5 Launches on Amazon Bedrock and Claude Platform on AWS

Anthropic's Claude Sonnet 5.5 is now available on Amazon Bedrock and Claude Platform on AWS, positioned as a faster, lower-cost model for well-scoped coding and document tasks. It pairs with the recently released Claude Opus 5.5, which handles higher-judgment work.

Comments

Loading...