model release

Alibaba Releases Qwen3.8 Max, a Multimodal Reasoning Model with 1M Token Context

TL;DR

Alibaba has moved Qwen3.8 Max out of preview into general availability, positioning it as the flagship of the Qwen3.8 series with a 1 million token context window and multimodal input support. The model is priced at $2.00 per million input tokens and $6.00 per million output tokens via OpenRouter.

2 min read
0

Alibaba's Qwen team has released Qwen3.8 Max, the general-availability successor to the Qwen3.8 Max Preview model. The release is now accessible through OpenRouter's API under the identifier qwen/qwen3.8-max.

Key Specifications

Qwen3.8 Max ships with a 1,000K (1 million) token context window, according to Alibaba. The model accepts text, image, and video inputs and produces text outputs, making it a multimodal reasoning model rather than a text-only system.

Pricing on OpenRouter is set at $2.00 per million input tokens and $6.00 per million output tokens.

Alibaba describes Qwen3.8 Max as the flagship model in the Qwen3.8 series, intended according to the company for complex reasoning and visual understanding tasks. The company has not disclosed the model's parameter count, training data cutoff date, or benchmark scores at time of publication.

What Changed From Preview

This release supersedes the Qwen3.8 Max Preview version. Alibaba has not published a detailed changelog specifying architectural or capability differences between the preview and general-availability builds. Without disclosed benchmark comparisons, the practical improvements over the preview remain according to Alibaba's release notes rather than independently verified.

Availability

The model is live now through OpenRouter, giving developers immediate API access without a separate preview or waitlist process. This mirrors Alibaba's pattern with prior Qwen releases, where preview models are used to gather feedback before a stable, generally available version ships with finalized pricing.

What this means

A 1 million token context window puts Qwen3.8 Max in the same tier as long-context leaders like Gemini's largest context configurations, at least on paper. Combined with multimodal input (text, image, video) and reasoning-oriented positioning, Alibaba is clearly targeting enterprise and developer use cases that require processing large documents, codebases, or video alongside complex reasoning chains.

The pricing — $2.00/M input and $6.00/M output — sits in a competitive but not aggressively cheap tier compared to other frontier-adjacent models available on OpenRouter. Without published benchmark scores (MMLU, GPQA, or coding evaluations), buyers have no independent way to verify Alibaba's reasoning and visual-understanding claims against competitors like GPT-4o, Claude, or Gemini 2.0 series models.

The bigger open question is whether the jump from "Preview" to general availability reflects meaningful capability gains or primarily a stability and SLA commitment for production use. Until Alibaba or third parties publish comparative benchmarks, developers evaluating Qwen3.8 Max should treat the reasoning and multimodal claims as unverified and test against their own workloads before committing.

Related Articles

model release

Alibaba Releases Qwen3.8 Max (0902), a 2.4-Trillion-Parameter MoE Model With 1M-Token Context

Alibaba's Qwen team released Qwen3.8 Max (0902), a 2.4-trillion-parameter mixture-of-experts model with a 1M-token context window that accepts text, image, and video input. The snapshot is post-trained for coding, agentic workflows, and long-horizon task execution, priced at $2/$6 per 1M input/output tokens.

model release

Meta Releases Muse Spark 1.3, a Free Multimodal Reasoning Model with 1M-Token Context

Meta has released Muse Spark 1.3, a multimodal reasoning model with a 1M-token context window, listed as free on OpenRouter. The model targets long-running agentic, multi-agent, and coding workflows, though audio input support remains incomplete.

model release

Meta Releases Muse Spark 1.3 Contributor, a Low-Cost Multimodal Reasoning Model With 1M Context Window

Meta has released Muse Spark 1.3 Contributor, described as the cost-efficient contributor tier of its multimodal reasoning model line. The model offers a 1 million token context window at $0.10 per 1M input tokens and $0.20 per 1M output tokens, targeting experimentation and early-stage agentic workflows.

model release

Google Lists Gemini 3.8 Flash on OpenRouter With 1M-Token Context, September 2026 Release Date

Google's Gemini 3.8 Flash has surfaced on OpenRouter with a 1-million-token context window and discounted pricing of $0.75 per 1M input tokens and $3.75 per 1M output tokens. Google has not issued a separate public announcement, and the listed release date of September 2, 2026 is unusually far out, leaving key details unconfirmed.

Comments

Loading...