Qwen3-Next-80B-A3B

Alibaba / Qwen🇨🇳 China
active

Architecture preview for Qwen3.5: hybrid attention (Gated DeltaNet + Gated Attention) with ultra-sparse MoE — 80B params, only 3B active. 10x cheaper to train and faster at long context than Qwen3-32B.

Context window262K tokens

Version History

qwen3-next-80b-a3b-launchmajor

Qwen3-Next-80B-A3B launches. Architecture preview for Qwen3.5: hybrid attention (Gated DeltaNet + Gated Attention) with ultra-sparse MoE — 80B params, only 3B active. 10x cheaper to train and faster at long context than Qwen3-32B.

Benchmark Scores

Full leaderboard →
9.3%
Hallucination Rate