Qwen3-Next-80B-A3B
Alibaba / Qwen🇨🇳 China
Architecture preview for Qwen3.5: hybrid attention (Gated DeltaNet + Gated Attention) with ultra-sparse MoE — 80B params, only 3B active. 10x cheaper to train and faster at long context than Qwen3-32B.
Context window262K tokens
Version History
qwen3-next-80b-a3b-launchmajor
Qwen3-Next-80B-A3B launches. Architecture preview for Qwen3.5: hybrid attention (Gated DeltaNet + Gated Attention) with ultra-sparse MoE — 80B params, only 3B active. 10x cheaper to train and faster at long context than Qwen3-32B.
Benchmark Scores
Full leaderboard →9.3%
Hallucination Rate