GLM-5.1 released: 754B agentic model outperforms Claude on coding benchmarks
Zhipu AI released GLM-5.1, a 754-parameter model optimized for agentic engineering tasks. The model scores 58.4% on SWE-Bench Pro, outperforming Claude 3.5 Sonnet (57.3%), and demonstrates sustained reasoning capability over hundreds of iterations.
GLM-5.1 — Quick Specs
GLM-5.1: 754B Agentic Model Outperforms Claude on Coding Benchmarks
Zhipu AI released GLM-5.1, a 754-parameter flagship model built for agentic engineering tasks. The model achieves 58.4% on SWE-Bench Pro—the primary metric for software engineering capability—exceeding Claude 3.5 Sonnet (57.3%) and maintaining substantial leads on specialized benchmarks.
Key Performance Metrics
GLM-5.1 achieves state-of-the-art performance across multiple agentic benchmarks:
- SWE-Bench Pro: 58.4% (vs. Claude 57.3%, Gemini 3.1 Pro 54.2%)
- NL2Repo (repo generation): 42.7% (vs. Claude 49.8%, significant improvement over GLM-5's 35.9%)
- Terminal-Bench 2.0: 63.5% on Terminus-2 suite
- CyberGym: 68.7% (vs. Claude 66.6%)
- BrowseComp with context management: 79.3% (vs. Gemini 84.0%, Claude 75.9%)
Mathematical reasoning shows mixed performance: 95.3% on AIME 2026 and 86.2% on GPQA-Diamond, trailing GPT-5.4 (98.7% on AIME) and Gemini 3.1 Pro (94.3% on GPQA).
Distinctive Agentic Capability
Unlike previous models including GLM-5, which plateau after initial optimizations, GLM-5.1 is designed to sustain effectiveness over extended problem-solving horizons. According to the developers, the model handles ambiguous problems with improved judgment and maintains productivity across longer sessions—breaking complex tasks into experiments, reading results, identifying blockers, and revising strategies through hundreds of iterations and thousands of tool calls.
This iterative reasoning approach distinguishes it from models optimized for single-pass performance.
Deployment and Quantization
Unsloth released GGUF quantized versions with 17 quant options ranging from 206 GB (1-bit UD-IQ1_M) to 1.51 TB (16-bit BF16). The releases implement Unsloth Dynamic 2.0 quantization, which the developers claim achieves superior accuracy compared to other quantization methods.
Supported inference frameworks include:
- SGLang (v0.5.10+)
- vLLM (v0.19.0+)
- xLLM (v0.8.0+)
- Transformers (v4.5.3+)
- KTransformers (v0.5.3+)
The model received 13,329 downloads on Hugging Face in its first month.
Availability
GLM-5.1 is available through Z.ai API Platform for inference. The developers announced chat.z.ai access would come in subsequent days. A technical report and GitHub repository were published alongside the release.
What This Means
GLM-5.1 represents a shift in agentic model design: instead of pursuing raw benchmark scores on isolated tasks, the focus is extended-horizon reasoning and iterative refinement. Its SWE-Bench Pro lead over Claude positions it as the strongest open-access model for software engineering tasks, though Gemini 3.1 Pro and GPT-5.4 maintain mathematical reasoning advantages. The quantized GGUF versions enable local deployment at scale, with memory requirements scaling from 206 GB to 1.51 TB depending on precision needs.
Related Articles
Alibaba Launches Wan3.0, Generating AI Videos Up to 30 Seconds From Text, Images, and Documents
Alibaba's Wan3.0 video generation model is now in beta, producing clips up to 30 seconds long and accepting text, images, video, audio, and documents like PDFs and PowerPoint files as input. Pricing runs per second of output across two tiers and three resolutions.
Mystery 'Stealth Model' Ox Alpha Appears on OpenRouter, Sparking Speculation Over Its Creator
A new AI model called Ox Alpha appeared on OpenRouter this week, listed only as being built by an anonymous third-party provider. Speculation about its creator has ranged from Z.ai's GLM models to an unreleased version of Microsoft's MAI, with no confirmation yet.
DeepSeek Releases Experimental V4-Flash-Vision-Exp, Claims Near-Parity With Opus 4.8 on Agent Benchmarks
DeepSeek has released V4-Flash-Vision-Exp, an experimental multimodal extension of V4-Flash that adds image understanding while preserving text reasoning capabilities. The company claims the model approaches or beats Anthropic's Opus 4.8 on its internal multimodal agent benchmarks.
DeepSeek Releases V4 Flash Vision Exp, an Experimental Multimodal MoE Model with 1M Context
DeepSeek has released V4 Flash Vision Exp, an experimental vision-enabled variant of DeepSeek V4 Flash 0731 that adds image understanding while matching the base model's text performance. The sparse mixture-of-experts model uses 13B active parameters out of 284B total and supports a 1M token context window.
Comments
Loading...