model release

Z.ai Launches GLM-5.3-Flash: 1M-Token Context, Image Support, Claimed 10x Cost Cut Over GLM-5.2

TL;DR

Z.ai has released GLM-5.3-Flash, a 320-billion-parameter Mixture-of-Experts model with 18 billion active parameters, a 1-million-token context window, and image input support. The model launched on LM Studio's Bionic platform hours after its official unveiling, with LM Studio claiming it is 9-10x cheaper to run than GLM-5.2.

2 min read
0

Z.ai has released GLM-5.3-Flash, a new Mixture-of-Experts model with a 1-million-token context window and native image input support. The model became available on LM Studio's Bionic agent platform within hours of its official announcement on August 26, 2026.

The specs

GLM-5.3-Flash is a 320-billion-parameter MoE model with 18 billion active parameters per forward pass, according to Z.ai. It accepts both image and text inputs and supports a 1-million-token context window. Before its official name was revealed, the model circulated under the codename "Ox Alpha" while being tested anonymously on OpenCode and OpenRouter, generating attention from developers who noticed its performance without knowing its origin.

Z.ai claims GLM-5.3-Flash surpasses its predecessor, GLM-5.2, across the benchmarks the company highlighted in its announcement, while generally landing in the same performance range as frontier models from Anthropic, OpenAI, Google, and DeepSeek. Z.ai did not publish specific benchmark scores alongside these claims, and independent verification is not yet available. Pricing has not been disclosed by Z.ai directly, though LM Studio states the model is 9-10x cheaper to run than GLM-5.2.

LM Studio's rapid integration

LM Studio added cloud support for GLM-5.3-Flash to Bionic, its Mac and Windows platform for agentic tasks like coding, research, and document handling. Bionic launched in mid-July 2026 and has since added models at a fast pace, including Moonshot AI's Kimi K3 shortly after launch.

Bionic supports both locally run models and cloud-hosted models. GLM-5.3-Flash is served from US-based servers with zero-data-retention (ZDR) enabled by default, matching the policy LM Studio applies to its other cloud offerings.

In a post on X, LM Studio wrote that GLM-5.3-Flash "surpasses GLM 5.2 in performance while being 9-10x cheaper" and confirmed image input support alongside the ZDR configuration.

What this means

The speed of this integration — cloud support live within hours of the model's public unveiling — signals how quickly platforms like LM Studio are moving to capture developer interest in new open or semi-open Chinese models. The reported 9-10x cost reduction from GLM-5.2 to GLM-5.3-Flash, if accurate, would make it one of the more aggressive price drops among MoE models this year, though neither Z.ai nor LM Studio has published exact per-token pricing to verify the multiple.

The 1-million-token context window and image support put GLM-5.3-Flash in direct competition with long-context multimodal models from major US labs, but until Z.ai releases specific benchmark numbers and independent evaluations emerge, claims that it matches frontier-model performance should be treated as unverified. The model's active-parameter count (18B against a 320B total) suggests Z.ai is optimizing for inference cost — consistent with the "Flash" branding and the claimed price advantage over GLM-5.2.

Related Articles

model release

Meta Releases Muse Spark 1.3, a Free Multimodal Reasoning Model with 1M-Token Context

Meta has released Muse Spark 1.3, a multimodal reasoning model with a 1M-token context window, listed as free on OpenRouter. The model targets long-running agentic, multi-agent, and coding workflows, though audio input support remains incomplete.

model release

InclusionAI Releases Ling 3.0 Flash Fin, a Finance-Focused MoE Model with 5.1B Active Parameters

InclusionAI has released Ling 3.0 Flash Fin, a finance-specialized mixture-of-experts model built on Ling 3.0 Flash. The model activates 5.1B of its 124B total parameters and targets long-horizon investment planning tasks while retaining general reasoning, coding, and math capabilities.

model release

Meta Releases Muse Spark 1.3 Contributor, a Low-Cost Multimodal Reasoning Model With 1M Context Window

Meta has released Muse Spark 1.3 Contributor, described as the cost-efficient contributor tier of its multimodal reasoning model line. The model offers a 1 million token context window at $0.10 per 1M input tokens and $0.20 per 1M output tokens, targeting experimentation and early-stage agentic workflows.

model release

Google Lists Gemini 3.8 Flash on OpenRouter With 1M-Token Context, September 2026 Release Date

Google's Gemini 3.8 Flash has surfaced on OpenRouter with a 1-million-token context window and discounted pricing of $0.75 per 1M input tokens and $3.75 per 1M output tokens. Google has not issued a separate public announcement, and the listed release date of September 2, 2026 is unusually far out, leaving key details unconfirmed.

Comments

Loading...