Alibaba Releases Qwen3.8-27B, a Dense Vision-Language Model with 1M-Token Context
Alibaba's Qwen team has released Qwen3.8-27B, a 27-billion-parameter dense vision-language model with 262,144-token native context extensible to 1 million tokens. The model shows gains over Qwen3.6-27B and Qwen3.7-Plus across coding, agentic, and multimodal benchmarks, according to Alibaba.
Qwen3.8-27B — Quick Specs
Qwen3.8-27B Launches with Native Vision-Language Support
Alibaba's Qwen team has released Qwen3.8-27B, a 27-billion-parameter dense causal language model with an integrated vision encoder, weights for which are now available on Hugging Face. The model is compatible with Hugging Face Transformers, vLLM, SGLang, and TokenSpeed, and is described by Alibaba as the most capable model yet in the open Qwen family.
Architecture and Specs
Qwen3.8-27B has 27 billion parameters, a hidden dimension of 5,120, and 64 layers arranged in a hybrid pattern of Gated DeltaNet and Gated Attention blocks alternating with feed-forward networks. It uses 48 linear attention heads for V and 16 for QK in its Gated DeltaNet layers, and 24 query heads with 4 key/value heads in its Gated Attention layers. The model was trained with multi-token prediction (MTP) across multiple steps.
Context length is 262,144 tokens natively, extensible up to 1,000,000 tokens, according to Alibaba. A hosted version via Qwen Cloud is planned with 1M-token context by default and built-in tools, though Alibaba says that service is "coming soon" with no pricing disclosed yet.
Benchmark Claims
According to Alibaba's published results, Qwen3.8-27B outperforms its predecessors Qwen3.6-27B and Qwen3.7-Plus across most tested categories. Reported scores include:
- Terminal-Bench 2.1 (Terminus): 73.0, versus 63.4 for Qwen3.6-27B and 64.0 for Qwen3.7-Plus
- SWE-bench Pro: 61.7, versus 53.5 and 57.6 respectively
- LiveCodeBench v6: 90.3, versus 83.9 and 89.6
- GPQA Diamond: 89.2, versus 87.8 and 90.3
- HLE (Humanity's Last Exam, judged by GPT-4o): 30.8, versus 24.0 and 34.7
- OSWorld-Verified (computer use): 84.3, versus 63.9 and 73.3
- AndroidWorld (mobile use): 81.9, versus 70.3 and 81.0
On several agentic and multimodal benchmarks, including WebArena-Verified, RecreationBench, and SWE-MM, Qwen3.8-27B posted the highest scores among the models compared, according to Alibaba's tables. It trailed the comparison model "Opus4.6 Max" on some benchmarks, including Terminal-Bench 2.1 (78.2) and HLE (40.0). All benchmark figures come from Alibaba and have not been independently verified.
Capabilities
Alibaba lists native image and video understanding, including hour-scale video and STEM diagrams, as a core feature. Thinking mode is enabled by default but can be disabled per request, with reasoning depth adjustable via a reasoning_effort parameter and prior reasoning retained through preserve_thinking. The company also emphasizes improvements in autonomous planning and handling of environment feedback for long-horizon agentic tasks.
What This Means
Qwen3.8-27B pushes Alibaba's open-model lineup toward larger native context windows and tighter integration of vision-language capability in a mid-sized, deployment-friendly dense model. The extensibility to 1 million tokens and strong reported scores on computer-use and coding-agent benchmarks position it as a candidate for agentic and multimodal workloads that previously required larger or proprietary models. Independent verification of the benchmark claims, and pricing for the eventual Qwen Cloud hosted version, remain outstanding.
Related Articles
Apple Releases LensVLM-9B, a 9B Vision-Language Model That Selectively Decompresses Text Images
Apple has released LensVLM-9B, a 9-billion-parameter vision-language model fine-tuned from Qwen3.5-9B-Base that processes documents as compressed images, selectively expanding only relevant pages to full resolution. The model supports 5x, 10x, and 15x compression ratios and is available under Apple's Machine Learning Research Model License.
Perceptron Launches Mk1.5, a Multimodal Perception Model for Physical Agents with Structured Spatial Outputs
Perceptron has released Mk1.5, a perception model built for physical agents that accepts text, image, video, and audio input and returns text alongside structured spatial annotations. It succeeds Mk1 and is priced at $0.15 per 1M input tokens and $1.50 per 1M output tokens.
Anonymous Stealth Model "Space Bunny Alpha" Debuts on OpenRouter With 1M-Token Context, Free During Preview
A previously unknown AI provider has released Space Bunny Alpha, a stealth model on OpenRouter offering a 1M-token context window, adjustable reasoning effort, and multimodal input support. The model is free during its preview period, though its developer remains unnamed.
OpenAI Scraps Release of GPT-6.1 Astra Over Safety Concerns
OpenAI confirmed it will not release GPT-6.1 Astra after the model failed to meet internal safety and alignment standards. The decision follows renewed industry-wide calls, including from Anthropic, to slow the pace of frontier model development.
Comments
Loading...