Z.ai Launches GLM-5.3-Flash With 1M-Token Context and Hybrid Attention Architecture
Z.ai has released GLM-5.3-Flash, a native multimodal model built for coding and long-horizon agent tasks, featuring a 1M-token context window and a hybrid sparse-linear attention architecture. The model is available via OpenRouter at a discounted $0.075/$0.25 per 1M tokens through September 2026.
GLM-5.3-Flash — Quick Specs
Z.ai has released GLM-5.3-Flash, a native multimodal model designed for efficient coding and long-horizon agent workloads. The model is now listed on OpenRouter with a 1 million token context window.
Architecture and Positioning
According to Z.ai, GLM-5.3-Flash uses a hybrid sparse and linear attention architecture intended to preserve accuracy over long-context inputs while cutting compute overhead relative to standard dense attention. The company positions the model for coding tasks and multi-step agent workflows that require sustained context tracking over extended sessions.
The model is described as "native multimodal," indicating it processes multiple input modalities directly rather than through bolted-on adapters, though Z.ai has not published a detailed breakdown of supported input types (text, image, audio, etc.) in the material reviewed.
Pricing and Availability
GLM-5.3-Flash is priced at:
- Input: $0.075 per 1M tokens (list price $0.15, currently 50% off)
- Output: $0.25 per 1M tokens (list price $0.50, currently 50% off)
- Cache read: $0.015–$0.03 per 1M tokens depending on provider
The 50% discount is available through the Z.ai provider on OpenRouter until September 9, 2026, at 16:00 UTC. Standard rates apply after that window closes.
The model officially released on August 26, 2026, and is currently served through OpenRouter with a reported P50 latency of 3.28 seconds and throughput of 27 tokens per second on the best-performing provider. OpenRouter lists 24-hour availability at 98.71% across the past three days of monitoring.
What We Don't Know
Z.ai has not disclosed a parameter count, training data cutoff date, or independent benchmark scores (MMLU, HumanEval, or similar) for GLM-5.3-Flash in the material reviewed. Claims about the hybrid attention architecture's compute savings and long-context accuracy come directly from Z.ai and have not been independently verified.
What This Means
GLM-5.3-Flash extends Z.ai's GLM line into a lower-cost, higher-throughput tier aimed squarely at coding assistants and agent frameworks that need to hold large amounts of context — codebases, tool logs, multi-turn plans — without paying dense-attention compute costs. At $0.075/$0.25 per 1M tokens, it undercuts many flagship multimodal models on price while matching or exceeding their context windows at 1M tokens.
The real test will be third-party benchmarking once independent evaluations surface, since Z.ai's efficiency and accuracy claims for the hybrid sparse-linear attention design are unverified. For now, the aggressive discount pricing and OpenRouter's multi-provider routing suggest Z.ai is pushing for rapid adoption among developers building coding agents, betting on volume before the promotional pricing window closes in September 2026.
Related Articles
Google releases Nano Banana 2.1 image model: $1.50/$30 per 1M tokens, 66K context
Google's Nano Banana 2.1 (Gemini Nano Banana 2.1) is an image generation and editing model on the Flash tier, listed on OpenRouter at $1.50 input and $30 output per 1M tokens with a 66K context window. It supports 1K, 2K, and 4K output and succeeds Nano Banana 2 and Nano Banana Pro, according to the listing.
Microsoft releases Decision-1, a Qwen3.5-9B-based model for classification and routing, at $0.042 per 1M input tokens
Microsoft has released Decision-1, a decision model built on Qwen3.5-9B for classification, evaluation, and routing. Microsoft claims 83.5% accuracy across 36 benchmarks and 85 ms latency. Input tokens cost $0.042 per million, and output tokens are free.
StepFun releases Step 5 Preview: 600B MoE with 1M context at $1/$2.70 per 1M tokens
StepFun has listed Step 5 Preview, a sparse Mixture-of-Experts model with 600B total and 27B active parameters and a 1.0M-token context window. It is priced at $1 input and $2.70 output per 1M tokens on OpenRouter. StepFun positions it as its flagship model for agentic work.
Liquid AI releases open d1-3B decision model: 16 ms on Jetson AGX Thor, 48.57 on Decision Index
Liquid AI released two open-weight decision models, d1-3B (text and image) and the experimental d1-omni-600M (text with image or audio). Unlike generative models, they answer in a single forward pass, and Liquid AI claims d1-3B scores 48.57 on its Decision Index 0.2.1, ahead of all 4B and 9B models it tested.
Comments
Loading...