Qwen Releases Qwen3.8 2.4T A95B, a 2.4-Trillion-Parameter Open-Weight MoE Model
Qwen has released Qwen3.8 2.4T A95B, an open-weight sparse mixture-of-experts model with 2.4 trillion total parameters and 95 billion active parameters per forward pass. The model is the open-weight variant of Qwen3.8 Max, targeting coding, research, complex reasoning, and agentic workflows with a 262K token context window.
Qwen Ships Open-Weight Variant of Qwen3.8 Max
Qwen has released Qwen3.8 2.4T A95B, a sparse mixture-of-experts (MoE) model with 2.4 trillion total parameters and 95 billion active parameters per inference pass. According to Qwen, this is the open-weight version of Qwen3.8 Max, the company's flagship closed model, and is designed for coding, research, complex reasoning, and agentic workflows.
The model is listed on OpenRouter with a 262,000-token context window and API pricing of $2 per 1M input tokens and $6 per 1M output tokens. The listing shows a release date of August 12, 2026.
Architecture
The headline figures — 2.4 trillion total parameters with only 95 billion active — put Qwen3.8 2.4T A95B among the largest MoE models disclosed to date by total parameter count, while keeping per-token compute closer to a mid-sized dense model through sparse expert routing. This design allows the model to draw on a much larger parameter pool for capacity while activating only a fraction of it (roughly 4%) for any given token, which is standard practice for MoE architectures used to balance quality against inference cost.
No benchmark scores, training data cutoff, or license details were included in the source listing. Qwen has not yet published a technical report, model card, or independent benchmark results for this release, so all architectural claims — including the total and active parameter counts — currently rest on Qwen's own listing rather than third-party verification.
What's confirmed vs. unconfirmed
Confirmed from the listing: total parameters (2.4T), active parameters (95B), context window (262K tokens), and API pricing ($2/$6 per 1M tokens). Unconfirmed: benchmark performance, training cutoff date, exact license terms, and whether the "open-weight" designation includes full weight downloads or a restricted release. Qwen describes this as the open-weight counterpart to a closed Qwen3.8 Max model, but no separate announcement for Qwen3.8 Max itself was available at time of writing.
What this means
A 2.4-trillion-parameter open-weight release, if the figures hold up under independent testing, would represent one of the largest total-parameter open models available for download and self-hosting, following the industry trend of decoupling total capacity from active compute via MoE routing. The lack of published benchmarks makes it impossible to assess how the model performs against comparable systems like DeepSeek V3, Llama 4, or Mistral's largest offerings. Practically, running a model with 2.4T total parameters — even with sparse activation — requires substantial memory and infrastructure for self-hosting, meaning most users will access it through API pricing at $2/$6 per 1M tokens rather than local deployment. Until Qwen publishes a technical report with verified benchmark scores and training details, the model's real-world capabilities relative to its scale remain an open question.
Related Articles
Alibaba Launches Qwen-Audio-3.1, Cuts AI Audio API Prices by Up to 95 Percent
Alibaba's Qwen team has released Qwen-Audio-3.1, a five-model lineup covering speech recognition, text-to-speech, and real-time voice interaction. Alongside the release, Alibaba cut API pricing by up to 95 percent for ASR, 85 percent for real-time models, and 70 percent for TTS.
Apple Releases LensVLM-9B, a 9B Vision-Language Model That Selectively Decompresses Text Images
Apple has released LensVLM-9B, a 9-billion-parameter vision-language model fine-tuned from Qwen3.5-9B-Base that processes documents as compressed images, selectively expanding only relevant pages to full resolution. The model supports 5x, 10x, and 15x compression ratios and is available under Apple's Machine Learning Research Model License.
Meta Releases Muse Glimmer 30B, an Open-Weight Agentic Model for Consumer Hardware
Meta Superintelligence Labs has released Muse Glimmer 30B, a dense open-weight model distilled from its larger Muse Spark system and tuned for agentic workflows on consumer hardware. The model supports 131K context, image understanding, and over 100 languages at $0.30/$1.10 per 1M input/output tokens.
NVIDIA Releases Nemotron 3 Diarization, an Open-Weight Speaker ID Model Supporting Up to 8 Speakers
NVIDIA has released Nemotron 3 Diarization, an open-weight speaker diarization model that determines "who spoke when" in audio, supporting both streaming and offline inference for up to eight speakers. The model achieves input buffer latency as low as 80 milliseconds and is available for commercial and non-commercial use.
Comments
Loading...