model releaseZhipu AI

Zhipu AI's GLM-5.1 outperforms GPT-5.4 and Claude Opus 4.6 on SWE-Bench Pro through iterative strategy refinement

TL;DR

Zhipu AI has released GLM-5.1, a freely available open-weight model designed for long-running programming tasks that achieves 58.4% on SWE-Bench Pro, edging out GPT-5.4 (57.7%) and Claude Opus 4.6 (57.3%). The model's core capability is iterative strategy refinement—it rethinks its approach across hundreds of iterations and thousands of tool calls, recognizing dead ends and shifting tactics without human intervention. However, GLM-5.1 trails on reasoning and knowledge benchmarks, scoring 31% on Humanity's Last Exam compared to Gemini 3.1 Pro's 45%.

3 min read
0

GLM-5.1 — Quick Specs

Context window203K tokens
Input$1.4/1M tokens
Output$4.4/1M tokens

Zhipu AI Releases GLM-5.1 with Iterative Long-Horizon Coding Capabilities

Zhipu AI has introduced GLM-5.1, an open-weight model available under MIT license on Hugging Face and ModelScope. The model is purpose-built for extended programming tasks where iterative strategy refinement, rather than raw parameter count, determines success on complex problems.

SWE-Bench Pro Leadership

On the SWE-Bench Pro software engineering benchmark, GLM-5.1 scores 58.4%—the highest among freely available models tested. This edges out:

  • GPT-5.4: 57.7%
  • Claude Opus 4.6: 57.3%

On CyberGym (cybersecurity), GLM-5.1 leads with 68.7%, though Zhipu AI notes that Gemini 3.1 Pro and GPT-5.4 refused some tasks for safety reasons, potentially affecting their scores.

Iterative Refinement Across Hundreds of Rounds

The model's defining feature is its ability to repeatedly review and revise its own strategy without external guidance. Zhipu AI demonstrates this through three internal evaluations:

Vector Database Optimization: GLM-5.1 improved query performance from Claude Opus 4.6's baseline of 3,547 queries per second to 21,500 queries per second—a 6.1x improvement. This required 600+ iterations and 6,000+ tool calls. The model initiated six major structural shifts:

  • Iteration 90: Switched from exhaustive search to clustering
  • Iteration 240: Introduced two-stage pipeline with pre-sorting and filtering

GPU Optimization: On KernelBench Level 3, GLM-5.1 achieved 3.6x speedup on baseline ML code versus Claude Opus 4.6's 4.2x. The model sustained progress longer than GLM-5 but remains behind the strongest competitor.

Linux Desktop Construction: When tasked with building a complete Linux desktop environment from a single prompt, GLM-5.1 delivered a functional system with file browser, terminal, text editor, system monitor, calculator, and games after eight hours of iterative refinement.

Mixed Results on Reasoning Tasks

GLM-5.1 shows clear weaknesses in non-coding domains:

  • Humanity's Last Exam (knowledge): 31% (vs. Gemini 3.1 Pro: 45%, GPT-5.4: 39.8%)
  • GPQA-Diamond (scientific reasoning): 86.2% (vs. Gemini 3.1 Pro: 94.3%, GPT-5.4: 92%)
  • Vending Bench 2 (agent business simulation): $5,634 balance (vs. Claude Opus 4.6: $8,018)
  • NL2Repo (repository generation): 42.7% (vs. Claude Opus 4.6: 49.8%)

On the Artificial Analysis Intelligence Index, GLM-5.1 ranks just behind Claude 4.6 Sonnet.

Acknowledged Limitations

Zhipu AI openly identifies remaining challenges: the model needs to recognize dead ends sooner, maintain coherence across thousands of tool calls, and reliably self-assess performance on tasks without clear success metrics. The company explicitly describes GLM-5.1 as a "first step."

Availability and Integration

GLM-5.1 is accessible via:

  • Hugging Face and ModelScope repositories
  • api.z.ai and BigModel.cn API platforms
  • Z.ai chat interface (launching in coming days)
  • Local deployment via vLLM and SGLang inference frameworks

The model integrates with coding agents including Claude Code and OpenClaw.

Market Context

GLM-5.1 represents Zhipu AI's expansion into autonomous coding—the company previously released GLM-5 (744B parameters) in February 2026 and GLM-5V-Turbo (multimodal coding) more recently. Competitors include Moonshot AI's Kimi K2.5 and Alibaba's Qwen3.5, both targeting the same agent-based coding market.

What This Means

GLM-5.1 demonstrates that extended iteration on coding tasks can outperform single-pass approaches from larger proprietary models—but only in specialized domains. The benchmarks reveal a critical trade-off: models optimized for iterative refinement on engineering problems lose generality on reasoning and knowledge tasks. The lack of independent verification for the three internal demonstrations leaves claims about iteration counts and strategy shifts unconfirmed. For coding-specific workloads, the open-weight availability and MIT license make GLM-5.1 worth evaluation; for general-purpose reasoning, leading proprietary models retain clear advantages.

Related Articles

model release

Mistral Large 4 enters public preview: 1T-parameter open-weight multimodal model, weights due by end of October

Mistral AI has launched a public preview of Mistral Large 4, a 1-trillion-parameter natively multimodal model with 49 billion active parameters. The preview API is live on Mistral Studio, and open weights are promised by the end of October 2026. Pricing and context window have not been disclosed.

model release

Reflection unveils 501B-parameter Beam, Mistral previews 1T-parameter Large 4, both open-weight

Reflection introduced Beam, a 501B-parameter mixture-of-experts model with 23B active parameters. Mistral said it is finishing Mistral Large 4, a 1T-parameter multimodal model with 49B active parameters. Both companies plan open-weight releases in October, and both are positioning the models against Chinese open-weight leaders.

model release

Reflection AI unveils Beam: 501B-parameter open-weight MoE with 1M-token context

Reflection AI has unveiled Beam, a text-only mixture-of-experts model with 501 billion total parameters, 23 billion active, and a 1 million token context window. The company claims it matches Z.ai's GLM-5.2 on advanced reasoning benchmarks while using 3-4x less inference compute. Weights and the full technical report are due later this month.

model release

StepFun releases Step 5 Preview: 600B MoE with 1M context at $1/$2.70 per 1M tokens

StepFun has listed Step 5 Preview, a sparse Mixture-of-Experts model with 600B total and 27B active parameters and a 1.0M-token context window. It is priced at $1 input and $2.70 output per 1M tokens on OpenRouter. StepFun positions it as its flagship model for agentic work.

Comments

Loading...