model release

Z.ai Releases GLM-5.3-Prime, a High-Throughput Variant of GLM-5.3 with 1M-Token Context

TL;DR

Z.ai has released GLM-5.3-Prime, a high-speed variant of its GLM-5.3 model that delivers 1.5-2x the output throughput through inference acceleration while retaining the full 1M-token context window. The model is priced at $2.80 per 1M input tokens and $8.80 per 1M output tokens, targeting coding and long-horizon agentic workloads.

2 min read
0

GLM-5.3-Prime — Quick Specs

Context window1000K tokens
Input$2.8/1M tokens
Output$8.8/1M tokens

Z.ai has released GLM-5.3-Prime, a high-speed inference variant of its GLM-5.3 model, according to a listing on OpenRouter. The company claims the model delivers 1.5–2x the output throughput of standard GLM-5.3 through inference acceleration, while inheriting the same full capabilities.

Specs and Pricing

GLM-5.3-Prime supports text input and output with a 1M-token context window and up to 128K output tokens. Pricing is set at $2.80 per 1M input tokens and $8.80 per 1M output tokens, with cached input reads priced at $0.56 per 1M tokens. The model became available via OpenRouter on September 23, 2026, served through Alibaba Cloud International infrastructure.

On that provider, OpenRouter reports a P50 latency of 0.70 seconds and throughput of roughly 45 tokens per second — figures that reflect current routing performance rather than a peak benchmark claim. OpenRouter's own 24-hour availability tracking shows 98.66% uptime for the model.

Reasoning and Target Use Cases

Reasoning is always enabled on GLM-5.3-Prime and cannot be turned off. Three reasoning-effort settings are available — low, high, and max — with max set as the default. Z.ai positions the model for coding and agentic workloads specifically, citing long-horizon multi-turn agent orchestration, real-time conversation, and streaming code generation as primary use cases.

This places GLM-5.3-Prime alongside a growing lineup of Z.ai models built around agentic and coding tasks, including the standard GLM-5.3 (priced at $0.45/$2.00 or $0.5614/$1.764 depending on provider), the multimodal GLM-5.3-Flash series, and earlier releases like GLM-5.2 and GLM-5.1. Z.ai has not disclosed parameter counts or independent benchmark scores for GLM-5.3-Prime; throughput and speed claims come directly from the company and have not been independently verified.

What This Means

GLM-5.3-Prime is not a new base model — it's a speed-optimized deployment of GLM-5.3's existing weights, aimed squarely at latency-sensitive agentic and coding pipelines where throughput matters as much as raw capability. The pricing sits above standard GLM-5.3 variants (which range from roughly $0.45 to $0.56 per 1M input tokens), suggesting Z.ai is charging a premium for the inference acceleration rather than offering it as a free upgrade.

For developers running long-horizon agent loops or streaming code-generation workloads, a 1.5–2x throughput gain — if it holds up under real-world load — could meaningfully cut wall-clock time on multi-step tasks without requiring a different model architecture. But because reasoning cannot be disabled and always defaults to maximum effort, users optimizing purely for cost or latency on simpler tasks may find the standard GLM-5.3 or the cheaper GLM-5.3-Flash line a better fit. The lack of independent benchmarks means the throughput and quality claims should be treated as unverified until third-party testing confirms them.

Comments

Loading...