Zhipu AI Releases GLM-5.3-Flash: First Multimodal Model in GLM-5 Series, 320B Parameters with Only 18B Active
Zhipu AI has released GLM-5.3-Flash, the first natively multimodal model in its GLM-5 series, built on a 320B-parameter mixture-of-experts architecture with only 18B active parameters. The company claims it outperforms GLM-5.2 across benchmarks at one-tenth the cost while approaching Claude Opus 4.8 on coding and agentic tasks.
Zhipu AI Ships GLM-5.3-Flash, a Multimodal MoE Model at 320B Parameters
Zhipu AI (zai-org) has released GLM-5.3-Flash, described as the first natively multimodal model in its GLM-5 series. The model uses a mixture-of-experts (MoE) design with 320 billion total parameters and 18 billion active parameters per forward pass, and weights are available on Hugging Face in BF16, F8_E4M3, and F32 tensor formats.
According to Zhipu AI, GLM-5.3-Flash outperforms its predecessor, GLM-5.2, across benchmarks and real-world workloads "at one-tenth the price," and approaches Claude Opus 4.8 on coding and agentic benchmarks. No specific benchmark score tables or dollar pricing figures were included in the release materials — these are company claims pending independent verification. Pricing is not yet disclosed.
Architecture Changes
GLM-5.3-Flash starts from a newly trained base model rather than a fine-tune of prior GLM checkpoints. Zhipu AI states the architecture introduces two notable changes for the GLM series:
- Hybrid sparse and linear attention, aimed at cutting long-context serving costs while preserving long-context accuracy.
- Manifold-Constrained Hyper-Connections (mHC), a technique the company says improves scaling efficiency during training.
The model was pre-trained on a 30-trillion-token multimodal corpus, larger and more varied than prior GLM training runs, which Zhipu AI credits for the efficiency gains.
Evaluation Notes
The release documentation lists evaluation protocols across several benchmarks, including HLE (with tools), NL2Repo, DeepSWE, Terminal-Bench 2.1, Agent's Last Exam (Toolathlon Verified), AutomationBench, GDPval-AA v2, and a vision test called BabyVision. Evaluations used context windows ranging up to 300,000 tokens (HLE) and up to 1 million tokens (NL2Repo), with some runs using GPT-5.6-luna as a judge model. No numeric scores for these benchmarks were published in the model card, so direct comparisons to GLM-5.2 or Claude Opus 4.8 cannot yet be independently confirmed.
Deployment
GLM-5.3-Flash supports local and self-hosted deployment through several inference frameworks: SGLang, vLLM, TokenSpeed, and KTransformers. It's also available via the Z.ai API Platform and through Hugging Face Inference Providers. The model card lists one fine-tune and two quantized variants already available on the Hugging Face model tree.
The work accompanies a technical report, "GLM-5: from Vibe Coding to Agentic Engineering," posted to arXiv (2602.15763) with a large multi-author list from the GLM-5 team, Zhipu AI, and Tsinghua University collaborators.
What This Means
GLM-5.3-Flash continues the trend of large MoE models activating a small fraction of total parameters to control inference cost — 18B active out of 320B total puts it in similar territory to other efficiency-focused MoE releases from Chinese labs this cycle. The addition of native multimodality to the GLM-5 line signals Zhipu AI is closing feature gaps with frontier multimodal models from OpenAI, Google DeepMind, and Anthropic.
The claims of matching Claude Opus 4.8 on coding and agentic benchmarks at a fraction of the cost are significant if verified, but the model card provides no concrete benchmark tables, pricing, or context window specification — only descriptions of evaluation methodology. Until third-party benchmark runs and official Z.ai API pricing are published, treat the performance and cost claims as unverified vendor statements rather than confirmed results.
Related Articles
Alibaba Releases Qwen3.8-Flash-Next: 125B-Parameter MoE Model Matches Larger Rivals at $0.16/$0.47 per Million Tokens
Alibaba's Qwen team released Qwen3.8-Flash-Next, a 125-billion-parameter mixture-of-experts model that activates just 6 billion parameters per token and previews architecture planned for Qwen4. The model outperforms the much larger Qwen3.7-Plus at roughly one-ninth the training cost and ships at $0.16 per million input tokens and $0.47 per million output tokens.
Z.ai Launches GLM-5.3-Flash With 1M-Token Context and Hybrid Attention Architecture
Z.ai has released GLM-5.3-Flash, a native multimodal model built for coding and long-horizon agent tasks, featuring a 1M-token context window and a hybrid sparse-linear attention architecture. The model is available via OpenRouter at a discounted $0.075/$0.25 per 1M tokens through September 2026.
Z.ai Releases GLM-5.3, Claims Frontier Coding Scores From a 750B-Parameter Model
Z.ai released GLM-5.3, a coding-focused model built on the same base as GLM-5.2 but with substantially extended post-training, and claims it surpasses Moonshot AI's Kimi K3 on many agentic coding benchmarks despite having roughly a third of the parameters. The model is live in Z.ai's coding plan now, with API and open-weight Hugging Face access expected within two weeks.
SenseNova Releases U1.5-8B-MoT, an Open-Weight Unified Model for Image Generation and Editing
SenseNova has released SenseNova-U1.5-8B-MoT, an open-weight native multimodal model built on its NEO-unify architecture for image generation, editing, and native 4K output. The model is available on Hugging Face under an Apache 2.0 license, with no inference pricing yet since it must be self-hosted.
Comments
Loading...