Alibaba Unveils Qwen3.8-Max, a 2.4T-Parameter Open-Weight Model for Coding and Agentic Work
Alibaba announced Qwen3.8-Max, a 2.4T-parameter flagship model targeting coding and long-horizon agentic work, with open weights promised for next week alongside Qwen3.8-27B. The model posted strong third-party benchmark results, ranking #4 in Frontend Code Arena and matching Claude Opus 4.7 on the Vals Index at roughly 2.3x lower cost.
Alibaba's Qwen Returns With a 2.4T-Parameter Flagship
Alibaba's Qwen team announced Qwen3.8-Max on August 3, 2026, a 2.4 trillion-parameter model the company calls its "most capable to date," with open weights promised for release the following week. A smaller Qwen3.8-27B model will also go open-weight, according to Alibaba.
The announcement matters because it ends speculation about whether Qwen would continue releasing competitive open models after last year's team restructuring, which had shifted focus toward closed API products. Qwen3.8-Max signals the lab is back in direct competition with other open-weight frontier releases, including Kimi K3 and GLM-5.2.
Pricing and Technical Specs
Alibaba priced Qwen3.8-Max API access at $2.00 per million input tokens, $6.00 per million output tokens, and $0.25 per million cached tokens. Third-party sources, including Vals AI and ZhihuFrontier, reported additional specs not confirmed directly by Alibaba: a 1 million-token context window, 128,000 max output tokens, approximately 95 billion active parameters per token (roughly 4% MoE activation ratio), and API support for low/medium/xhigh reasoning-effort modes with OpenAI and Anthropic protocol compatibility.
Benchmark Claims and Independent Evaluations
Alibaba claims the model can perform autonomous coding sessions lasting 10+ days, execute 500+ turns of chip design optimization, and run 365-day e-commerce strategy simulations. According to Alibaba, in one test the model achieved a 4.16x return (¥416,252 balance) in an e-commerce operations benchmark, and in another it improved a chip design from 8,298 to 678 gates while cutting die area by 81%.
Independent benchmarking largely corroborated strong performance, though not universally at the top:
- Frontend Code Arena: Qwen3.8-Max ranked #4 overall at 1,668 Elo, behind Claude Opus 5 [Max] (1,705) and Kimi K3 [Max] (1,676), according to Arena data.
- Vision Arena: #2 at 1,305, trailing Claude Fable 5 [High] by 13 points.
- Vals Index: #2 among open-weight models, #10 overall out of 43 tested, scoring 66.1 — matching Claude Opus 4.7's 66.1 score at roughly 2.3x lower cost ($2.68 vs. $6.17 per test run), according to Vals AI.
- SWE-bench: 87.3%, ahead of GPT-5.5 (82.6%) and GLM-5.2 (83.3%), but behind Claude Opus 4.8 (89.2%).
- Terminal-Bench 2.1: 67.4, up from 61.0 for the prior Qwen 3.7 Max release.
Vals AI noted the model gained 8.6 points on its Index score in roughly 2.5 months versus Qwen 3.7 Max (57.5 to 66.1), while API pricing dropped from $2.50/$7.50 to $2.00/$6.00 input/output over the same period.
Some claims remain anecdotal rather than benchmark-verified. One researcher on social media argued the model performs as the "best object detection VLM" across categories including satellite imagery and technical drawings, based on example outputs rather than a published benchmark. Another claimed the model surpassed Claude Fable 5 on Terminal-Bench specifically.
What This Means
Qwen3.8-Max's reported specs — a 2.4T-parameter count, 1M-token context, and MoE-style sparse activation — place it in the same deployment tier as other massive open-weight releases like Kimi K3 and GLM-5.2, rather than the more commonly self-hosted 30B–70B class. If the open-weight release next week matches the announced API performance, it would give developers access to a model competitive with Claude Opus-tier systems on coding benchmarks at a fraction of the cost — the Vals Index comparison shows roughly 2.3x lower per-test pricing versus Claude Opus 4.7 for matching scores.
The rapid benchmark gains between Qwen 3.7 Max and 3.8 Max (8.6 points in about 2.5 months) also suggest Alibaba is iterating faster than some closed-model competitors, adding pressure on labs like Anthropic and OpenAI in coding and agentic-workflow use cases specifically. However, until independent parties can run the open weights directly rather than relying on Alibaba's hosted API, some performance and safety characteristics remain unverified.
Related Articles
China Telecom Releases Xing4.0-29B-A4B, a 29B MoE Model Trained Entirely on Ascend NPUs
China Telecom Artificial Intelligence Technology has released Xing4.0-29B-A4B, a 29-billion-parameter mixture-of-experts model with only 4B parameters active per token and native 256K context. The company claims it is the first model of this scale trained entirely on Huawei's Ascend NPU platform using the MindSpore framework.
PrismML's Bonsai 2 Compresses 27B-Parameter Model to 5.9GB, Retains 98% of Benchmark Performance
PrismML released Bonsai 2 27B, a compressed version of Alibaba's Qwen3.8 27B model that shrinks memory footprint by 9x to 10x down to 5.9GB. The startup claims 98% aggregate benchmark parity with the original, up from 95% in its first release, using a ternary weight compression technique.
DeepSeek Ships V4.1-Flash With Novel Encoder-Decoder Architecture, Cuts KV Cache to 1/8 of Predecessor
DeepSeek released V4.1-Flash, a 763B-parameter model built on a new causal encoder-decoder architecture that splits 8B active parameters for prefill and 16B for decode. The model adds native vision support, a 1M-token context window, and shrinks KV cache footprint to roughly 1/8 of DeepSeek V4 Flash, while retiring V4 Pro.
InclusionAI Releases Ling 3.0 Flash VL, Adding Vision to Its 124B MoE Model
InclusionAI has released Ling 3.0 Flash VL, a vision-language extension of its 124B total-parameter, 5.5B active Mixture-of-Experts model. The model adds native image and video understanding, supports a 131K token context window, and is priced at $0.06 per 1M input tokens and $0.18 per 1M output tokens via OpenRouter.
Comments
Loading...