model release

Z.ai Releases GLM-5.2 with 1M Token Context Window at $1.40/$4.40 per Million

TL;DR

Z.ai has released GLM-5.2, a model designed for long-horizon engineering tasks with a 1 million token context window. The model is priced at $1.40 per million input tokens and $4.40 per million output tokens, and was released on June 16, 2025.

2 min read
0

GLM-5.2 — Quick Specs

Context window1000K tokens
Input$0.826/1M tokens
Output$2.596/1M tokens

Z.ai Releases GLM-5.2 with 1M Token Context Window

Z.ai has released GLM-5.2, a model with a 1 million token context window designed for project-level engineering tasks. The model is priced at $1.40 per million input tokens and $4.40 per million output tokens.

Technical Specifications

  • Context window: 1 million tokens
  • Input pricing: $1.40 per 1M tokens
  • Output pricing: $4.40 per 1M tokens
  • Release date: June 16, 2025
  • Modalities: Text in/out

Claimed Capabilities

According to Z.ai, GLM-5.2 is positioned as their "flagship model for the era of long-horizon tasks." The company claims the model can:

  • Handle project-level engineering context
  • Execute long-running tasks with improved reliability
  • Follow engineering standards consistently
  • Complete full development workflows from requirements to multi-platform deployment

The model is currently hosted exclusively through OpenRouter, with all requests forwarded directly to Z.ai's infrastructure.

Pricing Context

At $1.40/$4.40 per million tokens, GLM-5.2's pricing positions it in the mid-range for large context window models. For comparison:

  • Anthropic's Claude 3.5 Sonnet (200K context): $3/$15 per 1M tokens
  • OpenAI's GPT-4o (128K context): $2.50/$10 per 1M tokens
  • Google's Gemini 1.5 Pro (2M context): $1.25/$5 per 1M tokens

OpenRouter notes that effective pricing can be 60-80% lower than list prices when prompt caching is applied for repeated context.

What This Means

GLM-5.2 enters a competitive market for long-context models, where context window size alone no longer differentiates offerings. The 1M token window matches several existing models, while models like Gemini 1.5 Pro already offer 2M tokens at comparable pricing. The real test will be whether Z.ai's claimed advantages in long-horizon task execution and engineering workflow completion translate to measurable performance improvements in production use cases. Without published benchmark scores or independent verification of the model's capabilities, its market position remains uncertain.

Related Articles

model release

LG AI Research Releases K-EXAONE 2.0, a 750B-Parameter Open-Weight MoE Model with 262K Context

LG AI Research has released K-EXAONE 2.0, a 750-billion-parameter mixture-of-experts language model with 37B active parameters, a 262,144-token context window, and support for 10 languages. The model is open-weighted under Apache 2.0 and claims competitive results against Qwen3.5, GLM-5.1, and DeepSeek-V4 Pro on reasoning, coding, and long-context benchmarks.

model release

OpenAI Halts Parts of Astra Model Development After It Hit 'Critical' Cybersecurity Threshold

OpenAI disclosed that its in-development Astra model showed cyberattack capabilities strong enough that it cannot rule out a 'Critical' risk classification. The company has paused related internal activity and added security controls under its Preparedness Framework.

model release

Mistral's 3B-Parameter Shieldstral Matches 20B Safety Model on Text Benchmarks

Mistral's new Shieldstral, a 3-billion-parameter open-weight safety classifier, posts an 84.9% F1 score on text benchmarks—tying OpenAI's GPT-OSS-Safeguard-20B, a model roughly seven times larger. The model lets operators define safety rules at runtime using plain-language yes/no questions instead of fixed taxonomies.

model release

Mistral AI Releases Shieldstral-1.0-3B, a 3B-Parameter Policy-Adaptive Safety Classifier

Mistral AI has released Shieldstral-1.0-3B, a compact open-weight safety classifier that evaluates text and images against natural-language policies specified at inference time. The 3B model runs on a single GPU and reports F1 scores competitive with or exceeding larger moderation models like LlamaGuard-4-12B and GPT-OSS-Safeguard-20B on multiple benchmarks.

Comments

Loading...