model release

Z.ai Launches GLM-5.3-Flash With 1M-Token Context and Hybrid Attention Architecture

TL;DR

Z.ai has released GLM-5.3-Flash, a native multimodal model built for coding and long-horizon agent tasks, featuring a 1M-token context window and a hybrid sparse-linear attention architecture. The model is available via OpenRouter at a discounted $0.075/$0.25 per 1M tokens through September 2026.

2 min read
0

GLM-5.3-Flash — Quick Specs

Context window1000K tokens
Input$0.075/1M tokens
Output$0.25/1M tokens

Z.ai has released GLM-5.3-Flash, a native multimodal model designed for efficient coding and long-horizon agent workloads. The model is now listed on OpenRouter with a 1 million token context window.

Architecture and Positioning

According to Z.ai, GLM-5.3-Flash uses a hybrid sparse and linear attention architecture intended to preserve accuracy over long-context inputs while cutting compute overhead relative to standard dense attention. The company positions the model for coding tasks and multi-step agent workflows that require sustained context tracking over extended sessions.

The model is described as "native multimodal," indicating it processes multiple input modalities directly rather than through bolted-on adapters, though Z.ai has not published a detailed breakdown of supported input types (text, image, audio, etc.) in the material reviewed.

Pricing and Availability

GLM-5.3-Flash is priced at:

  • Input: $0.075 per 1M tokens (list price $0.15, currently 50% off)
  • Output: $0.25 per 1M tokens (list price $0.50, currently 50% off)
  • Cache read: $0.015–$0.03 per 1M tokens depending on provider

The 50% discount is available through the Z.ai provider on OpenRouter until September 9, 2026, at 16:00 UTC. Standard rates apply after that window closes.

The model officially released on August 26, 2026, and is currently served through OpenRouter with a reported P50 latency of 3.28 seconds and throughput of 27 tokens per second on the best-performing provider. OpenRouter lists 24-hour availability at 98.71% across the past three days of monitoring.

What We Don't Know

Z.ai has not disclosed a parameter count, training data cutoff date, or independent benchmark scores (MMLU, HumanEval, or similar) for GLM-5.3-Flash in the material reviewed. Claims about the hybrid attention architecture's compute savings and long-context accuracy come directly from Z.ai and have not been independently verified.

What This Means

GLM-5.3-Flash extends Z.ai's GLM line into a lower-cost, higher-throughput tier aimed squarely at coding assistants and agent frameworks that need to hold large amounts of context — codebases, tool logs, multi-turn plans — without paying dense-attention compute costs. At $0.075/$0.25 per 1M tokens, it undercuts many flagship multimodal models on price while matching or exceeding their context windows at 1M tokens.

The real test will be third-party benchmarking once independent evaluations surface, since Z.ai's efficiency and accuracy claims for the hybrid sparse-linear attention design are unverified. For now, the aggressive discount pricing and OpenRouter's multi-provider routing suggest Z.ai is pushing for rapid adoption among developers building coding agents, betting on volume before the promotional pricing window closes in September 2026.

Related Articles

model release

Zhipu AI Releases GLM-5.3-Flash: First Multimodal Model in GLM-5 Series, 320B Parameters with Only 18B Active

Zhipu AI has released GLM-5.3-Flash, the first natively multimodal model in its GLM-5 series, built on a 320B-parameter mixture-of-experts architecture with only 18B active parameters. The company claims it outperforms GLM-5.2 across benchmarks at one-tenth the cost while approaching Claude Opus 4.8 on coding and agentic tasks.

model release

Z.ai Confirmed as Creator of Chart-Topping 'Ox Alpha' Model, Weights Coming Wednesday

Z.ai, maker of the GLM model series, has confirmed it is behind Ox Alpha, the mysterious open-weight model that appeared anonymously on OpenRouter and topped benchmark leaderboards. The company will release the model's weights on Wednesday.

model release

Mystery 'Stealth Model' Ox Alpha Appears on OpenRouter, Sparking Speculation Over Its Creator

A new AI model called Ox Alpha appeared on OpenRouter this week, listed only as being built by an anonymous third-party provider. Speculation about its creator has ranged from Z.ai's GLM models to an unreleased version of Microsoft's MAI, with no confirmation yet.

model release

DeepSeek Releases Experimental V4-Flash-Vision-Exp, Claims Near-Parity With Opus 4.8 on Agent Benchmarks

DeepSeek has released V4-Flash-Vision-Exp, an experimental multimodal extension of V4-Flash that adds image understanding while preserving text reasoning capabilities. The company claims the model approaches or beats Anthropic's Opus 4.8 on its internal multimodal agent benchmarks.

Comments

Loading...