model release

Z.ai Releases GLM-5.2 with 1M Token Context Window at $1.40/$4.40 per Million

TL;DR

Z.ai has released GLM-5.2, a model designed for long-horizon engineering tasks with a 1 million token context window. The model is priced at $1.40 per million input tokens and $4.40 per million output tokens, and was released on June 16, 2025.

2 min read
0

GLM-5.2 — Quick Specs

Context window1000K tokens
Input$0.826/1M tokens
Output$2.596/1M tokens

Z.ai Releases GLM-5.2 with 1M Token Context Window

Z.ai has released GLM-5.2, a model with a 1 million token context window designed for project-level engineering tasks. The model is priced at $1.40 per million input tokens and $4.40 per million output tokens.

Technical Specifications

  • Context window: 1 million tokens
  • Input pricing: $1.40 per 1M tokens
  • Output pricing: $4.40 per 1M tokens
  • Release date: June 16, 2025
  • Modalities: Text in/out

Claimed Capabilities

According to Z.ai, GLM-5.2 is positioned as their "flagship model for the era of long-horizon tasks." The company claims the model can:

  • Handle project-level engineering context
  • Execute long-running tasks with improved reliability
  • Follow engineering standards consistently
  • Complete full development workflows from requirements to multi-platform deployment

The model is currently hosted exclusively through OpenRouter, with all requests forwarded directly to Z.ai's infrastructure.

Pricing Context

At $1.40/$4.40 per million tokens, GLM-5.2's pricing positions it in the mid-range for large context window models. For comparison:

  • Anthropic's Claude 3.5 Sonnet (200K context): $3/$15 per 1M tokens
  • OpenAI's GPT-4o (128K context): $2.50/$10 per 1M tokens
  • Google's Gemini 1.5 Pro (2M context): $1.25/$5 per 1M tokens

OpenRouter notes that effective pricing can be 60-80% lower than list prices when prompt caching is applied for repeated context.

What This Means

GLM-5.2 enters a competitive market for long-context models, where context window size alone no longer differentiates offerings. The 1M token window matches several existing models, while models like Gemini 1.5 Pro already offer 2M tokens at comparable pricing. The real test will be whether Z.ai's claimed advantages in long-horizon task execution and engineering workflow completion translate to measurable performance improvements in production use cases. Without published benchmark scores or independent verification of the model's capabilities, its market position remains uncertain.

Related Articles

model release

AllSpark's Iris-mini and Iris-pro Top Open-Weight Search Agent Benchmarks

Chinese lab AllSpark has released Iris-mini and Iris-pro, two open-weight search agents built on Qwen3 models that claim the top spot among open-weight systems in their size classes on four research benchmarks. The release includes model weights, an agent harness, and evaluation code, with training pipelines to follow.

model release

Tencent Open-Sources AuK, a 1.5B-Parameter Speech Generation and Editing Model

Tencent has open-sourced AuK, a 1.5B-parameter foundation model for speech generation and editing that handles TTS, content editing, and audio enhancement through natural-language instructions. The release includes a distilled AuK-Flash variant for 4-step fast inference, both under MIT license.

model release

Google Releases TimesFM-3, a 330M-Parameter Model That Forecasts Sales Using Weather and Discount Data

Google Research has released TimesFM-3, a 330-million-parameter time series forecasting model that predicts outcomes like sales by combining related variables, historical data, and known future events such as discounts or weather. The model claims top rankings on three benchmarks against Amazon's Chronos-2 and the Toto-2.0 family.

model release

DeepSeek Ships V4.1-Flash With Novel Encoder-Decoder Architecture, Cuts KV Cache to 1/8 of Predecessor

DeepSeek released V4.1-Flash, a 763B-parameter model built on a new causal encoder-decoder architecture that splits 8B active parameters for prefill and 16B for decode. The model adds native vision support, a 1M-token context window, and shrinks KV cache footprint to roughly 1/8 of DeepSeek V4 Flash, while retiring V4 Pro.

Comments

Loading...