Alibaba Releases Qwen3.8 Max (0902), a 2.4-Trillion-Parameter MoE Model With 1M-Token Context
Alibaba's Qwen team released Qwen3.8 Max (0902), a 2.4-trillion-parameter mixture-of-experts model with a 1M-token context window that accepts text, image, and video input. The snapshot is post-trained for coding, agentic workflows, and long-horizon task execution, priced at $2/$6 per 1M input/output tokens.
Alibaba Ships Qwen3.8 Max (0902) Snapshot
Alibaba's Qwen team has released Qwen3.8 Max (0902), an updated snapshot of its Qwen3.8 Max model. According to the model listing on OpenRouter, it is a 2.4-trillion-parameter mixture-of-experts (MoE) model with a 1-million-token context window that accepts text, image, and video input and returns text output.
The model has reasoning enabled by default, with configurable reasoning effort exposed through the API. Alibaba states the snapshot was post-trained specifically for coding and agentic work, including multi-step software projects, multi-tool orchestration, and long-horizon task execution. It also targets chart reasoning, document parsing, and multimodal understanding across long documents and extended video.
Pricing and Access
The model is priced at $2 per 1M input tokens and $6 per 1M output tokens through Alibaba Cloud Intelligence, the listed provider on OpenRouter. Cache read pricing is listed at $0.25 per 1M tokens, with 5-minute cache creation at $2.50 per 1M tokens and 5-minute cache read at $0.17 per 1M tokens. The listing shows an availability window of 3 days at the time of publication, with uptime data not yet available.
Capabilities
According to Alibaba, Qwen3.8 Max (0902) supports:
- Tool calling and structured outputs for agentic workflows
- Multi-step, multi-tool task orchestration
- Long-horizon task execution for extended software projects
- Chart reasoning and document parsing
- Multimodal understanding over long documents and extended video
- Configurable reasoning effort
No independent benchmark scores were included in the source listing, so performance claims on coding, agentic, or multimodal tasks remain unverified pending third-party evaluation.
What This Means
A 2.4-trillion-parameter MoE model with a 1M-token context window puts Qwen3.8 Max (0902) in the same scale tier as the largest frontier models currently shipping from US labs, at least on paper. The pricing — $2/$6 per 1M tokens — undercuts many Western frontier-model equivalents, continuing Alibaba's pattern of aggressive pricing on high-capability checkpoints.
The emphasis on agentic and coding post-training, combined with default-on reasoning and tool calling, signals Qwen is positioning this snapshot for developers building autonomous coding agents and long-running multi-step workflows rather than general chat use cases. The lack of published benchmark scores in this listing means buyers will need to run their own evaluations before committing production traffic. As with prior Qwen snapshots, expect independent benchmark comparisons against GPT and Claude model families to surface within days of wider availability.
Related Articles
Meta Releases Muse Spark 1.3, a Free Multimodal Reasoning Model with 1M-Token Context
Meta has released Muse Spark 1.3, a multimodal reasoning model with a 1M-token context window, listed as free on OpenRouter. The model targets long-running agentic, multi-agent, and coding workflows, though audio input support remains incomplete.
OpenAI's GPT-6 Astra Reportedly Automates AI Engineering Tasks at Under $6 an Hour, According to Latent Space Testing
A Latent Space report describes GPT-6 Astra, a new OpenAI model the blog says can autonomously handle AI engineering tasks—training models, labeling data, deploying systems—at an estimated cost of under $6 per hour. The claims, including 97.6% on FrontierMath and 99.9% on ARC-AGI-3, come from independent blog testing rather than an official OpenAI announcement.
InclusionAI Releases Ling 3.0 Flash Fin, a Finance-Focused MoE Model with 5.1B Active Parameters
InclusionAI has released Ling 3.0 Flash Fin, a finance-specialized mixture-of-experts model built on Ling 3.0 Flash. The model activates 5.1B of its 124B total parameters and targets long-horizon investment planning tasks while retaining general reasoning, coding, and math capabilities.
Meta Releases Muse Spark 1.3 Contributor, a Low-Cost Multimodal Reasoning Model With 1M Context Window
Meta has released Muse Spark 1.3 Contributor, described as the cost-efficient contributor tier of its multimodal reasoning model line. The model offers a 1 million token context window at $0.10 per 1M input tokens and $0.20 per 1M output tokens, targeting experimentation and early-stage agentic workflows.
Comments
Loading...