Alibaba Releases Qwen3.8 Open-Weight Models Under Apache 2.0, Including 27B Multimodal Model with 262K Native Context
Alibaba's Qwen team has released open weights for Qwen3.8, including a 27-billion-parameter multimodal dense model with 262,000 tokens of native context. The models ship under the Apache 2.0 license and are available on Hugging Face and ModelScope.
Alibaba's Qwen team has released open model weights for Qwen3.8, a new family of models that includes a 27-billion-parameter multimodal dense model and a significantly larger mixture-of-experts model. Both ship under the Apache 2.0 license, allowing unrestricted commercial use.
The flagship model, Qwen3.8-27B, is a dense multimodal model with 27 billion parameters. According to Qwen, it outperforms the larger Qwen3.7-Plus model on coding and office tasks, though this claim has not been independently verified with published benchmark scores. The team also states the model shows improved agent capabilities, planning tasks more independently and completing them more reliably than prior Qwen releases.
Qwen3.8-27B natively handles up to 262,000 tokens of context and can scale to 1 million tokens using the YaRN extrapolation method. Beyond text, the model processes images and video, including diagrams, documents, and video content running several hours long. A thinking mode — Qwen's term for extended chain-of-thought reasoning before responding — is enabled by default but can be toggled off per query.
Alongside the 27B model, Qwen released weights for a much larger model, Qwen3.8-2.4T-A95B, designed to operate at what the company calls its "Max" performance tier. The naming convention indicates a mixture-of-experts architecture with 2.4 trillion total parameters and roughly 95 billion active parameters per forward pass, consistent with Qwen's prior naming scheme for large MoE models.
Both models are available now on Hugging Face and ModelScope. Alibaba says a hosted version with the full 1 million-token context window will be available soon through Qwen Cloud, the company's AI service platform. Pricing for the Qwen Cloud hosted version has not yet been disclosed.
No independent benchmark scores, training data cutoff date, or third-party evaluations were provided alongside the release. Claims of outperforming Qwen3.7-Plus and improved agentic behavior currently rest solely on statements from the Qwen team.
What this means
Qwen continues its aggressive release cadence for open-weight models, following a pattern of shipping both efficient dense models and massive MoE models under permissive licensing. The 262K native context window — extendable to 1 million tokens — puts Qwen3.8-27B in direct competition with long-context leaders like Gemini and Claude, but at a fraction of the parameter count if the performance claims hold up.
The simultaneous release of a 27B dense model and a 2.4T-parameter MoE model signals Alibaba's strategy of covering both edge/on-premise deployment and cloud-scale "Max" use cases from a single model family. Because the weights are open under Apache 2.0, independent labs and developers can now verify Qwen's coding and agentic performance claims directly — something the AI community has increasingly demanded as vendor benchmark claims proliferate without third-party confirmation. Until those independent evaluations arrive, the coding and agent capability claims should be treated as unverified.
Related Articles
Alibaba Releases Qwen3.8-2.4T-A95B-FP8: 2.4T-Parameter Open Model with 1M-Token Context
Alibaba's Qwen team has released Qwen3.8-2.4T-A95B-FP8, an open-weight, FP8-quantized MoE model with 2.4 trillion total parameters and 95 billion activated per token. It natively supports 262,144 tokens of context, extensible to 1,010,000, and forms the base for the hosted Qwen3.8-Max API.
Alibaba Releases Qwen3.8-27B-FP8, a 27B Dense Vision-Language Model with 1M-Token Context
Alibaba's Qwen team has released FP8-quantized weights for Qwen3.8-27B, a 27-billion-parameter dense vision-language model with native 262,144-token context extensible to 1 million tokens. The model claims gains over its Qwen3.6 and Qwen3.7 predecessors on coding, agentic, and multimodal benchmarks.
Alibaba Releases Qwen3.8-27B, a Dense Vision-Language Model with 1M-Token Context
Alibaba's Qwen team has released Qwen3.8-27B, a 27-billion-parameter dense vision-language model with 262,144-token native context extensible to 1 million tokens. The model shows gains over Qwen3.6-27B and Qwen3.7-Plus across coding, agentic, and multimodal benchmarks, according to Alibaba.
Qwen Releases Qwen3.8 2.4T A95B, a 2.4-Trillion-Parameter Open-Weight MoE Model
Qwen has released Qwen3.8 2.4T A95B, an open-weight sparse mixture-of-experts model with 2.4 trillion total parameters and 95 billion active parameters per forward pass. The model is the open-weight variant of Qwen3.8 Max, targeting coding, research, complex reasoning, and agentic workflows with a 262K token context window.
Comments
Loading...