model release

Z.ai Launches GLM-5.3-Flash: 1M-Token Context, Image Support, Claimed 10x Cost Cut Over GLM-5.2

TL;DR

Z.ai has released GLM-5.3-Flash, a 320-billion-parameter Mixture-of-Experts model with 18 billion active parameters, a 1-million-token context window, and image input support. The model launched on LM Studio's Bionic platform hours after its official unveiling, with LM Studio claiming it is 9-10x cheaper to run than GLM-5.2.

2 min read
0

Z.ai has released GLM-5.3-Flash, a new Mixture-of-Experts model with a 1-million-token context window and native image input support. The model became available on LM Studio's Bionic agent platform within hours of its official announcement on August 26, 2026.

The specs

GLM-5.3-Flash is a 320-billion-parameter MoE model with 18 billion active parameters per forward pass, according to Z.ai. It accepts both image and text inputs and supports a 1-million-token context window. Before its official name was revealed, the model circulated under the codename "Ox Alpha" while being tested anonymously on OpenCode and OpenRouter, generating attention from developers who noticed its performance without knowing its origin.

Z.ai claims GLM-5.3-Flash surpasses its predecessor, GLM-5.2, across the benchmarks the company highlighted in its announcement, while generally landing in the same performance range as frontier models from Anthropic, OpenAI, Google, and DeepSeek. Z.ai did not publish specific benchmark scores alongside these claims, and independent verification is not yet available. Pricing has not been disclosed by Z.ai directly, though LM Studio states the model is 9-10x cheaper to run than GLM-5.2.

LM Studio's rapid integration

LM Studio added cloud support for GLM-5.3-Flash to Bionic, its Mac and Windows platform for agentic tasks like coding, research, and document handling. Bionic launched in mid-July 2026 and has since added models at a fast pace, including Moonshot AI's Kimi K3 shortly after launch.

Bionic supports both locally run models and cloud-hosted models. GLM-5.3-Flash is served from US-based servers with zero-data-retention (ZDR) enabled by default, matching the policy LM Studio applies to its other cloud offerings.

In a post on X, LM Studio wrote that GLM-5.3-Flash "surpasses GLM 5.2 in performance while being 9-10x cheaper" and confirmed image input support alongside the ZDR configuration.

What this means

The speed of this integration — cloud support live within hours of the model's public unveiling — signals how quickly platforms like LM Studio are moving to capture developer interest in new open or semi-open Chinese models. The reported 9-10x cost reduction from GLM-5.2 to GLM-5.3-Flash, if accurate, would make it one of the more aggressive price drops among MoE models this year, though neither Z.ai nor LM Studio has published exact per-token pricing to verify the multiple.

The 1-million-token context window and image support put GLM-5.3-Flash in direct competition with long-context multimodal models from major US labs, but until Z.ai releases specific benchmark numbers and independent evaluations emerge, claims that it matches frontier-model performance should be treated as unverified. The model's active-parameter count (18B against a 320B total) suggests Z.ai is optimizing for inference cost — consistent with the "Flash" branding and the claimed price advantage over GLM-5.2.

Related Articles

model release

Z.ai Launches GLM-5.3-Flash With 1M-Token Context and Hybrid Attention Architecture

Z.ai has released GLM-5.3-Flash, a native multimodal model built for coding and long-horizon agent tasks, featuring a 1M-token context window and a hybrid sparse-linear attention architecture. The model is available via OpenRouter at a discounted $0.075/$0.25 per 1M tokens through September 2026.

model release

Zhipu AI Releases GLM-5.3-Flash: First Multimodal Model in GLM-5 Series, 320B Parameters with Only 18B Active

Zhipu AI has released GLM-5.3-Flash, the first natively multimodal model in its GLM-5 series, built on a 320B-parameter mixture-of-experts architecture with only 18B active parameters. The company claims it outperforms GLM-5.2 across benchmarks at one-tenth the cost while approaching Claude Opus 4.8 on coding and agentic tasks.

model release

Alibaba Releases Qwen3.8 Flash, a Multimodal Reasoning Model with 1M-Token Context

Alibaba has released Qwen3.8 Flash, a multimodal reasoning model with a 1 million token context window, aimed at coding, agentic workflows, and visual/document analysis. It's priced at $0.16 per 1M input tokens and $0.47 per 1M output tokens through Alibaba Cloud International.

model release

DeepSeek Releases V4 Flash Vision Exp, an Experimental Multimodal MoE Model with 1M Context

DeepSeek has released V4 Flash Vision Exp, an experimental vision-enabled variant of DeepSeek V4 Flash 0731 that adds image understanding while matching the base model's text performance. The sparse mixture-of-experts model uses 13B active parameters out of 284B total and supports a 1M token context window.

Comments

Loading...