GLM-5.3-Flash

Zhipu AI🇨🇳 China
active
Context window1000K tokens
Input / 1M tokens$0.075
Output / 1M tokens$0.25

Version History

5.3-flashminor

Z.ai released GLM-5.3-Flash, a native multimodal model with a 1M-token context window and a hybrid sparse-linear attention architecture aimed at coding and agent workloads. It launched with discounted pricing of $0.075/$0.25 per 1M input/output tokens through September 2026.

Benchmark Scores

Full leaderboard →
1475.0 elo
Arena Elo
91.2%
GPQA
53.0 tokens_per_sec
Speed (tok/s)

Coverage