Qwen3.8-Flash-Next

Alibaba / Qwen🇨🇳 China
active
Context window262K tokens

Version History

3.8-Flash-Nextmajor

Qwen3.8-Flash-Next is an experimental preview of the architecture planned for Qwen4, introducing Qwen Sparse Attention, gated residual streams, n-gram embeddings, and a new optimizer recipe. It has 125B total parameters with 6B activated and supports up to 1 million tokens of context.

Benchmark Scores

Full leaderboard →
91.7%
GPQA
53.0 tokens_per_sec
Speed (tok/s)

Coverage