Alibaba Unveils Qwen3.8-Max, a 2.4T-Parameter Open-Weight Model for Coding and Agentic Work
Alibaba announced Qwen3.8-Max, a 2.4T-parameter flagship model targeting coding and long-horizon agentic work, with open weights promised for next week alongside Qwen3.8-27B. The model posted strong third-party benchmark results, ranking #4 in Frontend Code Arena and matching Claude Opus 4.7 on the Vals Index at roughly 2.3x lower cost.
Alibaba's Qwen Returns With a 2.4T-Parameter Flagship
Alibaba's Qwen team announced Qwen3.8-Max on August 3, 2026, a 2.4 trillion-parameter model the company calls its "most capable to date," with open weights promised for release the following week. A smaller Qwen3.8-27B model will also go open-weight, according to Alibaba.
The announcement matters because it ends speculation about whether Qwen would continue releasing competitive open models after last year's team restructuring, which had shifted focus toward closed API products. Qwen3.8-Max signals the lab is back in direct competition with other open-weight frontier releases, including Kimi K3 and GLM-5.2.
Pricing and Technical Specs
Alibaba priced Qwen3.8-Max API access at $2.00 per million input tokens, $6.00 per million output tokens, and $0.25 per million cached tokens. Third-party sources, including Vals AI and ZhihuFrontier, reported additional specs not confirmed directly by Alibaba: a 1 million-token context window, 128,000 max output tokens, approximately 95 billion active parameters per token (roughly 4% MoE activation ratio), and API support for low/medium/xhigh reasoning-effort modes with OpenAI and Anthropic protocol compatibility.
Benchmark Claims and Independent Evaluations
Alibaba claims the model can perform autonomous coding sessions lasting 10+ days, execute 500+ turns of chip design optimization, and run 365-day e-commerce strategy simulations. According to Alibaba, in one test the model achieved a 4.16x return (¥416,252 balance) in an e-commerce operations benchmark, and in another it improved a chip design from 8,298 to 678 gates while cutting die area by 81%.
Independent benchmarking largely corroborated strong performance, though not universally at the top:
- Frontend Code Arena: Qwen3.8-Max ranked #4 overall at 1,668 Elo, behind Claude Opus 5 [Max] (1,705) and Kimi K3 [Max] (1,676), according to Arena data.
- Vision Arena: #2 at 1,305, trailing Claude Fable 5 [High] by 13 points.
- Vals Index: #2 among open-weight models, #10 overall out of 43 tested, scoring 66.1 — matching Claude Opus 4.7's 66.1 score at roughly 2.3x lower cost ($2.68 vs. $6.17 per test run), according to Vals AI.
- SWE-bench: 87.3%, ahead of GPT-5.5 (82.6%) and GLM-5.2 (83.3%), but behind Claude Opus 4.8 (89.2%).
- Terminal-Bench 2.1: 67.4, up from 61.0 for the prior Qwen 3.7 Max release.
Vals AI noted the model gained 8.6 points on its Index score in roughly 2.5 months versus Qwen 3.7 Max (57.5 to 66.1), while API pricing dropped from $2.50/$7.50 to $2.00/$6.00 input/output over the same period.
Some claims remain anecdotal rather than benchmark-verified. One researcher on social media argued the model performs as the "best object detection VLM" across categories including satellite imagery and technical drawings, based on example outputs rather than a published benchmark. Another claimed the model surpassed Claude Fable 5 on Terminal-Bench specifically.
What This Means
Qwen3.8-Max's reported specs — a 2.4T-parameter count, 1M-token context, and MoE-style sparse activation — place it in the same deployment tier as other massive open-weight releases like Kimi K3 and GLM-5.2, rather than the more commonly self-hosted 30B–70B class. If the open-weight release next week matches the announced API performance, it would give developers access to a model competitive with Claude Opus-tier systems on coding benchmarks at a fraction of the cost — the Vals Index comparison shows roughly 2.3x lower per-test pricing versus Claude Opus 4.7 for matching scores.
The rapid benchmark gains between Qwen 3.7 Max and 3.8 Max (8.6 points in about 2.5 months) also suggest Alibaba is iterating faster than some closed-model competitors, adding pressure on labs like Anthropic and OpenAI in coding and agentic-workflow use cases specifically. However, until independent parties can run the open weights directly rather than relying on Alibaba's hosted API, some performance and safety characteristics remain unverified.
Related Articles
Alibaba Releases Qwen3.8-Max, a 2.4 Trillion-Parameter Model Built for Multi-Day Autonomous Tasks
Alibaba has released Qwen3.8-Max, a 2.4-trillion-parameter model with 95 billion active parameters per query, designed to run autonomous tasks over multiple days. The company claims it hits 93 on PaperBench and rivals Claude Opus 4.8 and GPT-5.6 Sol on internal benchmarks, with open weights arriving next week.
Alibaba Markets Qwen 3.8 as a Job Enhancer, Not a Job Killer — But Skips the Technical Specs
Alibaba is promoting its new Qwen 3.8 model with marketing that frames AI automation as liberating rather than threatening, a departure from the fear-based messaging common among Western AI labs. The company has not disclosed technical specifications, benchmark scores, or pricing for the model.
Alibaba Releases Qwen3.8 Max, a Multimodal Reasoning Model with 1M Token Context
Alibaba has moved Qwen3.8 Max out of preview into general availability, positioning it as the flagship of the Qwen3.8 series with a 1 million token context window and multimodal input support. The model is priced at $2.00 per million input tokens and $6.00 per million output tokens via OpenRouter.
OpenAI Reportedly Developing 'Astra' Model Family for Multi-Day Autonomous Problem-Solving
OpenAI is reportedly developing a new model family called Astra, designed to coordinate multiple agents on complex problems over hours or days. The models are already in testing and would be first to go through a planned U.S. government pre-release review, according to The Information.
Comments
Loading...