Alibaba Releases Qwen3.8-27B, a Dense Vision-Language Model with 1M-Token Context
Alibaba's Qwen team has released Qwen3.8-27B, a 27-billion-parameter dense vision-language model with 262,144-token native context extensible to 1 million tokens. The model shows gains over Qwen3.6-27B and Qwen3.7-Plus across coding, agentic, and multimodal benchmarks, according to Alibaba.
Qwen3.8-27B Launches with Native Vision-Language Support
Alibaba's Qwen team has released Qwen3.8-27B, a 27-billion-parameter dense causal language model with an integrated vision encoder, weights for which are now available on Hugging Face. The model is compatible with Hugging Face Transformers, vLLM, SGLang, and TokenSpeed, and is described by Alibaba as the most capable model yet in the open Qwen family.
Architecture and Specs
Qwen3.8-27B has 27 billion parameters, a hidden dimension of 5,120, and 64 layers arranged in a hybrid pattern of Gated DeltaNet and Gated Attention blocks alternating with feed-forward networks. It uses 48 linear attention heads for V and 16 for QK in its Gated DeltaNet layers, and 24 query heads with 4 key/value heads in its Gated Attention layers. The model was trained with multi-token prediction (MTP) across multiple steps.
Context length is 262,144 tokens natively, extensible up to 1,000,000 tokens, according to Alibaba. A hosted version via Qwen Cloud is planned with 1M-token context by default and built-in tools, though Alibaba says that service is "coming soon" with no pricing disclosed yet.
Benchmark Claims
According to Alibaba's published results, Qwen3.8-27B outperforms its predecessors Qwen3.6-27B and Qwen3.7-Plus across most tested categories. Reported scores include:
- Terminal-Bench 2.1 (Terminus): 73.0, versus 63.4 for Qwen3.6-27B and 64.0 for Qwen3.7-Plus
- SWE-bench Pro: 61.7, versus 53.5 and 57.6 respectively
- LiveCodeBench v6: 90.3, versus 83.9 and 89.6
- GPQA Diamond: 89.2, versus 87.8 and 90.3
- HLE (Humanity's Last Exam, judged by GPT-4o): 30.8, versus 24.0 and 34.7
- OSWorld-Verified (computer use): 84.3, versus 63.9 and 73.3
- AndroidWorld (mobile use): 81.9, versus 70.3 and 81.0
On several agentic and multimodal benchmarks, including WebArena-Verified, RecreationBench, and SWE-MM, Qwen3.8-27B posted the highest scores among the models compared, according to Alibaba's tables. It trailed the comparison model "Opus4.6 Max" on some benchmarks, including Terminal-Bench 2.1 (78.2) and HLE (40.0). All benchmark figures come from Alibaba and have not been independently verified.
Capabilities
Alibaba lists native image and video understanding, including hour-scale video and STEM diagrams, as a core feature. Thinking mode is enabled by default but can be disabled per request, with reasoning depth adjustable via a reasoning_effort parameter and prior reasoning retained through preserve_thinking. The company also emphasizes improvements in autonomous planning and handling of environment feedback for long-horizon agentic tasks.
What This Means
Qwen3.8-27B pushes Alibaba's open-model lineup toward larger native context windows and tighter integration of vision-language capability in a mid-sized, deployment-friendly dense model. The extensibility to 1 million tokens and strong reported scores on computer-use and coding-agent benchmarks position it as a candidate for agentic and multimodal workloads that previously required larger or proprietary models. Independent verification of the benchmark claims, and pricing for the eventual Qwen Cloud hosted version, remain outstanding.
Related Articles
Alibaba Releases Qwen3.8-27B-FP8, a 27B Dense Vision-Language Model with 1M-Token Context
Alibaba's Qwen team has released FP8-quantized weights for Qwen3.8-27B, a 27-billion-parameter dense vision-language model with native 262,144-token context extensible to 1 million tokens. The model claims gains over its Qwen3.6 and Qwen3.7 predecessors on coding, agentic, and multimodal benchmarks.
Alibaba Releases Qwen3.8-2.4T-A95B-FP8: 2.4T-Parameter Open Model with 1M-Token Context
Alibaba's Qwen team has released Qwen3.8-2.4T-A95B-FP8, an open-weight, FP8-quantized MoE model with 2.4 trillion total parameters and 95 billion activated per token. It natively supports 262,144 tokens of context, extensible to 1,010,000, and forms the base for the hosted Qwen3.8-Max API.
Qwen Releases Qwen3.8 2.4T A95B, a 2.4-Trillion-Parameter Open-Weight MoE Model
Qwen has released Qwen3.8 2.4T A95B, an open-weight sparse mixture-of-experts model with 2.4 trillion total parameters and 95 billion active parameters per forward pass. The model is the open-weight variant of Qwen3.8 Max, targeting coding, research, complex reasoning, and agentic workflows with a 262K token context window.
ByteDance Seed Launches Seed 2.1 Turbo, a 262K-Context Multimodal Model for Coding Agents
ByteDance Seed has released Seed 2.1 Turbo, a multimodal model targeting coding and long-horizon agent workflows with a 262K token context window. The model is priced at $0.50 per 1M input tokens and $2.50 per 1M output tokens, and is now listed on OpenRouter.
Comments
Loading...