GLM-5.3-FlashX

Zhipu AI🇨🇳 China
active
Context window1000K tokens

Version History

5.3-FlashXminor

Z.ai released GLM-5.3-FlashX as a high-speed variant of GLM-5.3-Flash, using a 320B total / 18B active parameter hybrid attention architecture to deliver up to 200 tokens/second inference with a 1M-token context window.

Coverage