Z.ai Releases GLM-5.2 with 1M Token Context Window at $1.40/$4.40 per Million
Z.ai has released GLM-5.2, a model designed for long-horizon engineering tasks with a 1 million token context window. The model is priced at $1.40 per million input tokens and $4.40 per million output tokens, and was released on June 16, 2025.
GLM-5.2 — Quick Specs
Z.ai Releases GLM-5.2 with 1M Token Context Window
Z.ai has released GLM-5.2, a model with a 1 million token context window designed for project-level engineering tasks. The model is priced at $1.40 per million input tokens and $4.40 per million output tokens.
Technical Specifications
- Context window: 1 million tokens
- Input pricing: $1.40 per 1M tokens
- Output pricing: $4.40 per 1M tokens
- Release date: June 16, 2025
- Modalities: Text in/out
Claimed Capabilities
According to Z.ai, GLM-5.2 is positioned as their "flagship model for the era of long-horizon tasks." The company claims the model can:
- Handle project-level engineering context
- Execute long-running tasks with improved reliability
- Follow engineering standards consistently
- Complete full development workflows from requirements to multi-platform deployment
The model is currently hosted exclusively through OpenRouter, with all requests forwarded directly to Z.ai's infrastructure.
Pricing Context
At $1.40/$4.40 per million tokens, GLM-5.2's pricing positions it in the mid-range for large context window models. For comparison:
- Anthropic's Claude 3.5 Sonnet (200K context): $3/$15 per 1M tokens
- OpenAI's GPT-4o (128K context): $2.50/$10 per 1M tokens
- Google's Gemini 1.5 Pro (2M context): $1.25/$5 per 1M tokens
OpenRouter notes that effective pricing can be 60-80% lower than list prices when prompt caching is applied for repeated context.
What This Means
GLM-5.2 enters a competitive market for long-context models, where context window size alone no longer differentiates offerings. The 1M token window matches several existing models, while models like Gemini 1.5 Pro already offer 2M tokens at comparable pricing. The real test will be whether Z.ai's claimed advantages in long-horizon task execution and engineering workflow completion translate to measurable performance improvements in production use cases. Without published benchmark scores or independent verification of the model's capabilities, its market position remains uncertain.
Related Articles
AllSpark's Iris-mini and Iris-pro Top Open-Weight Search Agent Benchmarks
Chinese lab AllSpark has released Iris-mini and Iris-pro, two open-weight search agents built on Qwen3 models that claim the top spot among open-weight systems in their size classes on four research benchmarks. The release includes model weights, an agent harness, and evaluation code, with training pipelines to follow.
Tencent Open-Sources AuK, a 1.5B-Parameter Speech Generation and Editing Model
Tencent has open-sourced AuK, a 1.5B-parameter foundation model for speech generation and editing that handles TTS, content editing, and audio enhancement through natural-language instructions. The release includes a distilled AuK-Flash variant for 4-step fast inference, both under MIT license.
Google Releases TimesFM-3, a 330M-Parameter Model That Forecasts Sales Using Weather and Discount Data
Google Research has released TimesFM-3, a 330-million-parameter time series forecasting model that predicts outcomes like sales by combining related variables, historical data, and known future events such as discounts or weather. The model claims top rankings on three benchmarks against Amazon's Chronos-2 and the Toto-2.0 family.
DeepSeek Ships V4.1-Flash With Novel Encoder-Decoder Architecture, Cuts KV Cache to 1/8 of Predecessor
DeepSeek released V4.1-Flash, a 763B-parameter model built on a new causal encoder-decoder architecture that splits 8B active parameters for prefill and 16B for decode. The model adds native vision support, a 1M-token context window, and shrinks KV cache footprint to roughly 1/8 of DeepSeek V4 Flash, while retiring V4 Pro.
Comments
Loading...