model release

Z.ai Releases GLM-5.3-FlashX, a 200 Tokens/Second Variant of Its GLM-5.3-Flash Model

TL;DR

Z.ai has released GLM-5.3-FlashX, a high-speed variant of GLM-5.3-Flash built on a hybrid sparse and linear attention architecture with 320B total parameters (18B active). The model supports a 1M-token context window and claims inference speeds of up to 200 tokens per second.

2 min read
0

Z.ai released GLM-5.3-FlashX on September 18, 2026, a high-speed variant of its GLM-5.3-Flash model designed for fast inference in coding, visual understanding, and long-horizon agent workloads.

Key Specifications

GLM-5.3-FlashX is a native multimodal model built on the same hybrid sparse and linear attention architecture that underlies GLM-5.3-Flash. According to Z.ai, the model uses 320B total parameters with 18B active parameters at inference time — a mixture-of-experts-style design intended to cut compute overhead while preserving accuracy over long context windows.

The model supports a 1M-token context window and, according to Z.ai, delivers inference speeds of up to 200 tokens per second. Pricing for GLM-5.3-FlashX has not yet been disclosed; the base GLM-5.3-Flash model it derives from is listed at $0.075 per 1M input tokens and $0.25 per 1M output tokens on OpenRouter, though FlashX's own rate has not been separately confirmed.

Positioning Within the GLM Lineup

GLM-5.3-FlashX sits within Z.ai's broader GLM-5.3 family, which also includes GLM-5.3-Flash (offered in both 1.3M and 1.0M context variants at $0.075/$0.25 per 1M tokens) and GLM-5.3, a reasoning model priced at $0.70-$0.90 per 1M input tokens and $2.20-$2.83 per 1M output tokens depending on context configuration. Unlike GLM-5.3, which keeps reasoning always on, GLM-5.3-FlashX is positioned as a speed-optimized, non-reasoning variant aimed at latency-sensitive coding and agent applications.

Z.ai has been iterating rapidly through its GLM series, moving from GLM-4.5V and GLM-4.6 through GLM-5, GLM-5.1, GLM-5.2, and now GLM-5.3 variants within roughly a year. The FlashX designation follows a pattern seen elsewhere in the industry, where companies ship a lower-latency, cost-reduced sibling alongside a flagship reasoning model to serve high-throughput production use cases.

What This Means

GLM-5.3-FlashX targets a specific gap: applications that need near-instant responses at large context sizes without paying the latency or cost premium of a full reasoning model. The 200 tokens/second claim, if it holds up in independent testing, would make it competitive for real-time coding assistants and agent loops that call models repeatedly within a single task.

The missing piece is pricing. Without a confirmed rate for FlashX specifically, it's difficult to assess whether the speed gains come with a cost trade-off relative to the standard GLM-5.3-Flash tier. The 320B/18B active parameter split suggests Z.ai is using aggressive sparsity to keep serving costs down even at high throughput, which is consistent with its stated goal of balancing performance against token efficiency across the GLM-5.3 generation. Independent benchmarks on coding and agentic tasks will be needed to verify Z.ai's throughput and capability claims.

Related Articles

model release

Unverified 'GPT Astra' Model Appears on OpenRouter With 1.05M Token Context, No OpenAI Confirmation

OpenRouter is listing a model called 'OpenAI GPT Astra Latest' with a 1.05 million token context window and $10/$50 per-million-token pricing. OpenAI has made no public announcement, and the listing's own description says it is an auto-redirecting alias rather than a fixed model.

model release

OpenRouter Lists 'GPT Sol Latest' — An Alias Pointer to OpenAI's Newest Sol-Family Model, Not a Standalone Release

OpenRouter has added a listing called '~openai/gpt-sol-latest,' described as an alias that always points to the newest model in an undisclosed 'GPT Sol' family from OpenAI. The listing shows a 1050K token context window and pricing of $2.00 per million input tokens and $10.00 per million output tokens, but OpenAI has not publicly confirmed a model line by this name.

model release

Google DeepMind Launches Gemini 3.8 Live, Claims #1 Spot on Speech-to-Speech Benchmark

Google DeepMind has released Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking, two voice-dialogue models that reason and execute background tasks without interrupting conversation. Google claims the Extended Thinking model ranks #1 on Artificial Analysis' Speech to Speech Quality Index with a score of 82.6.

model release

Moonshot AI's 2.8 Trillion-Parameter Kimi K3 Launches on Amazon Bedrock with 1M-Token Context

Moonshot AI's Kimi K3, described by the company as the first open model to reach 2.8 trillion parameters, is now available on Amazon Bedrock. It features native vision, a 1-million-token context window, and is the first open-weight model on Bedrock to support explicit prompt caching.

Comments

Loading...