Z.ai Launches GLM-5.3-Flash With 1M-Token Context and Hybrid Attention Architecture
Z.ai has released GLM-5.3-Flash, a native multimodal model built for coding and long-horizon agent tasks, featuring a 1M-token context window and a hybrid sparse-linear attention architecture. The model is available via OpenRouter at a discounted $0.075/$0.25 per 1M tokens through September 2026.
GLM-5.3-Flash — Quick Specs
Z.ai has released GLM-5.3-Flash, a native multimodal model designed for efficient coding and long-horizon agent workloads. The model is now listed on OpenRouter with a 1 million token context window.
Architecture and Positioning
According to Z.ai, GLM-5.3-Flash uses a hybrid sparse and linear attention architecture intended to preserve accuracy over long-context inputs while cutting compute overhead relative to standard dense attention. The company positions the model for coding tasks and multi-step agent workflows that require sustained context tracking over extended sessions.
The model is described as "native multimodal," indicating it processes multiple input modalities directly rather than through bolted-on adapters, though Z.ai has not published a detailed breakdown of supported input types (text, image, audio, etc.) in the material reviewed.
Pricing and Availability
GLM-5.3-Flash is priced at:
- Input: $0.075 per 1M tokens (list price $0.15, currently 50% off)
- Output: $0.25 per 1M tokens (list price $0.50, currently 50% off)
- Cache read: $0.015–$0.03 per 1M tokens depending on provider
The 50% discount is available through the Z.ai provider on OpenRouter until September 9, 2026, at 16:00 UTC. Standard rates apply after that window closes.
The model officially released on August 26, 2026, and is currently served through OpenRouter with a reported P50 latency of 3.28 seconds and throughput of 27 tokens per second on the best-performing provider. OpenRouter lists 24-hour availability at 98.71% across the past three days of monitoring.
What We Don't Know
Z.ai has not disclosed a parameter count, training data cutoff date, or independent benchmark scores (MMLU, HumanEval, or similar) for GLM-5.3-Flash in the material reviewed. Claims about the hybrid attention architecture's compute savings and long-context accuracy come directly from Z.ai and have not been independently verified.
What This Means
GLM-5.3-Flash extends Z.ai's GLM line into a lower-cost, higher-throughput tier aimed squarely at coding assistants and agent frameworks that need to hold large amounts of context — codebases, tool logs, multi-turn plans — without paying dense-attention compute costs. At $0.075/$0.25 per 1M tokens, it undercuts many flagship multimodal models on price while matching or exceeding their context windows at 1M tokens.
The real test will be third-party benchmarking once independent evaluations surface, since Z.ai's efficiency and accuracy claims for the hybrid sparse-linear attention design are unverified. For now, the aggressive discount pricing and OpenRouter's multi-provider routing suggest Z.ai is pushing for rapid adoption among developers building coding agents, betting on volume before the promotional pricing window closes in September 2026.
Related Articles
Unverified 'GPT Astra' Model Appears on OpenRouter With 1.05M Token Context, No OpenAI Confirmation
OpenRouter is listing a model called 'OpenAI GPT Astra Latest' with a 1.05 million token context window and $10/$50 per-million-token pricing. OpenAI has made no public announcement, and the listing's own description says it is an auto-redirecting alias rather than a fixed model.
OpenRouter Lists 'GPT Sol Latest' — An Alias Pointer to OpenAI's Newest Sol-Family Model, Not a Standalone Release
OpenRouter has added a listing called '~openai/gpt-sol-latest,' described as an alias that always points to the newest model in an undisclosed 'GPT Sol' family from OpenAI. The listing shows a 1050K token context window and pricing of $2.00 per million input tokens and $10.00 per million output tokens, but OpenAI has not publicly confirmed a model line by this name.
Inference.net Launches Schematron V2 Turbo, a 3B-Parameter Model for High-Volume HTML-to-JSON Extraction
Inference.net has released Schematron V2 Turbo, a 3-billion-parameter model built specifically for high-volume HTML-to-JSON extraction. The model supports a 128K context window and is priced at $0.03 per 1M input tokens and $0.15 per 1M output tokens.
Inference.net Releases Schematron V2 Small, a 3B-Parameter Model for HTML-to-JSON Extraction
Inference.net has released Schematron V2 Small, a 3B-parameter model specialized in converting HTML pages into structured JSON output. The model supports a 128K context window and requires extraction schemas to be passed via response_format rather than standard prompts.
Comments
Loading...