Alibaba Releases Qwen3.8-Flash-Next, a 125B-Parameter Preview of Qwen4's Architecture
Alibaba's Qwen team has released Qwen3.8-Flash-Next, an open-weight model with 125B total parameters (6B activated) that previews architectural changes planned for Qwen4, including a new sparse attention mechanism and n-gram embeddings. The model natively supports 262,144 tokens of context, extensible to 1 million.
Qwen3.8-Flash-Next — Quick Specs
Alibaba's Qwen team has released Qwen3.8-Flash-Next, an open-weight model described as an experimental preview of the architecture that will underpin the forthcoming Qwen4. The release introduces four architectural changes the company says are aimed at improving efficiency as models scale, rather than simply adding parameters or context length.
Architecture and Specs
Qwen3.8-Flash-Next has 125B total parameters with 6B activated per forward pass, plus a separate 51B-parameter n-gram embedding table and a 4B multi-token prediction (MTP) module. The model uses 48 layers with a hybrid layout: three Gated DeltaNet-plus-MoE blocks followed by one Qwen Sparse Attention (QSA)-plus-MoE block, repeated 12 times. Its Mixture-of-Experts layer routes to 10 of 512 experts plus one shared expert per token.
Four changes distinguish this architecture from prior Qwen releases, according to Alibaba:
- Hybrid Attention with QSA: Replaces the earlier Gated Attention pairing with Qwen Sparse Attention, which operates on micro-blocks rather than individual tokens, aimed at cutting long-context latency for agentic workloads.
- Gated Residual: Adds element-wise data-dependent read gates and per-branch scalar write gates to residual streams, intended to improve expressiveness without destabilizing training.
- N-gram Embedding: A 20-million-entry bigram/trigram embedding table (51B parameters) that the company says scales more cheaply than adding MoE capacity and is more suitable for offloading on memory-constrained hardware.
- Tailored Training Recipe: Applies Muon and AdamW optimizers to different weight categories and removes traditional batch-size warmup, which Alibaba claims reduces optimizer steps while allowing larger learning rates.
Context length is 262,144 tokens natively, extensible to 1,000,000 tokens. A separate production variant, Qwen3.8-Flash, ships with 1M context by default and built-in tools, according to Alibaba; this article covers the base Flash-Next checkpoint.
Benchmark Claims
Alibaba reports Qwen3.8-Flash-Next scoring 62.5 on SWE-bench Pro, versus 61.7 for Qwen3.8-27B and 56.0 for DeepSeek-V4-Flash-0731. On GPQA Diamond it scores 91.7, close to Claude-Opus-4.6 (Max)'s 91.3. On Humanity's Last Exam (HLE), it scores 35.9, below Claude-Opus-4.6's reported 40.0. On LiveCodeBench v6, the model scores 91.9, ahead of all listed comparisons including Claude-Opus-4.6 (Max) at 88.8. In vision-language tests, the company reports 84.5 on AndroidWorld and 19.4/52.3 (binary/partial) on OSWorld 2.0. These figures come from Alibaba's own technical report and have not been independently verified.
What this means
This release is not a general-availability flagship model — Alibaba explicitly frames it as an architectural preview, and pricing for API access has not been disclosed. The significance lies in the design choices: sparse attention operating at block granularity, n-gram embeddings as an alternative parameter-scaling axis to MoE, and a training recipe that drops batch-size warmup. If these hold up under independent testing, they suggest Alibaba is prioritizing inference efficiency and memory-constrained deployment over raw parameter growth for Qwen4. Until third parties reproduce the benchmark numbers, particularly the SWE-bench Pro and HLE results, they should be treated as vendor-reported figures rather than settled facts.
Related Articles
Qwen releases Qwen-Image-2.1-Turbo: 8-step text-to-image and editing checkpoint on a 7B architecture
Qwen has published Qwen-Image-2.1-Turbo on Hugging Face, an accelerated checkpoint of Qwen-Image-2.1 that runs text-to-image generation and image editing in 8 denoising steps. It keeps the same 7B visual generation architecture and loads through a new QwenImage21Pipeline in Diffusers.
StepFun releases Step 5 Preview: 600B MoE with 1M context at $1/$2.70 per 1M tokens
StepFun has listed Step 5 Preview, a sparse Mixture-of-Experts model with 600B total and 27B active parameters and a 1.0M-token context window. It is priced at $1 input and $2.70 output per 1M tokens on OpenRouter. StepFun positions it as its flagship model for agentic work.
Mistral Large 4 enters public preview: 1T-parameter open-weight multimodal model, weights due by end of October
Mistral AI has launched a public preview of Mistral Large 4, a 1-trillion-parameter natively multimodal model with 49 billion active parameters. The preview API is live on Mistral Studio, and open weights are promised by the end of October 2026. Pricing and context window have not been disclosed.
Reflection unveils 501B-parameter Beam, Mistral previews 1T-parameter Large 4, both open-weight
Reflection introduced Beam, a 501B-parameter mixture-of-experts model with 23B active parameters. Mistral said it is finishing Mistral Large 4, a 1T-parameter multimodal model with 49B active parameters. Both companies plan open-weight releases in October, and both are positioning the models against Chinese open-weight leaders.
Comments
Loading...