model release

Alibaba Releases Qwen3.8 Max, a Multimodal Reasoning Model with 1M Token Context

TL;DR

Alibaba has moved Qwen3.8 Max out of preview into general availability, positioning it as the flagship of the Qwen3.8 series with a 1 million token context window and multimodal input support. The model is priced at $2.00 per million input tokens and $6.00 per million output tokens via OpenRouter.

2 min read
0

Alibaba's Qwen team has released Qwen3.8 Max, the general-availability successor to the Qwen3.8 Max Preview model. The release is now accessible through OpenRouter's API under the identifier qwen/qwen3.8-max.

Key Specifications

Qwen3.8 Max ships with a 1,000K (1 million) token context window, according to Alibaba. The model accepts text, image, and video inputs and produces text outputs, making it a multimodal reasoning model rather than a text-only system.

Pricing on OpenRouter is set at $2.00 per million input tokens and $6.00 per million output tokens.

Alibaba describes Qwen3.8 Max as the flagship model in the Qwen3.8 series, intended according to the company for complex reasoning and visual understanding tasks. The company has not disclosed the model's parameter count, training data cutoff date, or benchmark scores at time of publication.

What Changed From Preview

This release supersedes the Qwen3.8 Max Preview version. Alibaba has not published a detailed changelog specifying architectural or capability differences between the preview and general-availability builds. Without disclosed benchmark comparisons, the practical improvements over the preview remain according to Alibaba's release notes rather than independently verified.

Availability

The model is live now through OpenRouter, giving developers immediate API access without a separate preview or waitlist process. This mirrors Alibaba's pattern with prior Qwen releases, where preview models are used to gather feedback before a stable, generally available version ships with finalized pricing.

What this means

A 1 million token context window puts Qwen3.8 Max in the same tier as long-context leaders like Gemini's largest context configurations, at least on paper. Combined with multimodal input (text, image, video) and reasoning-oriented positioning, Alibaba is clearly targeting enterprise and developer use cases that require processing large documents, codebases, or video alongside complex reasoning chains.

The pricing — $2.00/M input and $6.00/M output — sits in a competitive but not aggressively cheap tier compared to other frontier-adjacent models available on OpenRouter. Without published benchmark scores (MMLU, GPQA, or coding evaluations), buyers have no independent way to verify Alibaba's reasoning and visual-understanding claims against competitors like GPT-4o, Claude, or Gemini 2.0 series models.

The bigger open question is whether the jump from "Preview" to general availability reflects meaningful capability gains or primarily a stability and SLA commitment for production use. Until Alibaba or third parties publish comparative benchmarks, developers evaluating Qwen3.8 Max should treat the reasoning and multimodal claims as unverified and test against their own workloads before committing.

Related Articles

model release

Alibaba Markets Qwen 3.8 as a Job Enhancer, Not a Job Killer — But Skips the Technical Specs

Alibaba is promoting its new Qwen 3.8 model with marketing that frames AI automation as liberating rather than threatening, a departure from the fear-based messaging common among Western AI labs. The company has not disclosed technical specifications, benchmark scores, or pricing for the model.

model release

Alibaba Releases Qwen3.8-Max, a 2.4 Trillion-Parameter Model Built for Multi-Day Autonomous Tasks

Alibaba has released Qwen3.8-Max, a 2.4-trillion-parameter model with 95 billion active parameters per query, designed to run autonomous tasks over multiple days. The company claims it hits 93 on PaperBench and rivals Claude Opus 4.8 and GPT-5.6 Sol on internal benchmarks, with open weights arriving next week.

model release

Thinking Machines Lab Releases Inkling Small: 276B MoE Model with 524K Context Window

Thinking Machines Lab has released Inkling Small, an open-weight multimodal mixture-of-experts model with 12B active parameters out of 276B total and a 524K token context window. The model targets reasoning, coding, agentic workflows, and multilingual use cases at $0.58 per 1M input tokens and $1.44 per 1M output tokens.

model release

DeepSeek Releases V4 Flash 0731: 1M-Token MoE Model at $0.14/M Input Tokens

DeepSeek has released V4 Flash 0731, a sparse mixture-of-experts model with 13B active parameters out of 284B total and a 1049K token context window. The model targets coding, reasoning, and agent workflows, priced at $0.14 per million input tokens and $0.28 per million output tokens via OpenRouter.

Comments

Loading...