model release

Alibaba Releases Qwen3.8 Flash, a Multimodal Reasoning Model with 1M-Token Context

TL;DR

Alibaba has released Qwen3.8 Flash, a multimodal reasoning model with a 1 million token context window, aimed at coding, agentic workflows, and visual/document analysis. It's priced at $0.16 per 1M input tokens and $0.47 per 1M output tokens through Alibaba Cloud International.

2 min read
0

Qwen3.8 Flash — Quick Specs

Context window1000K tokens
Input$0.16/1M tokens
Output$0.47/1M tokens

Alibaba's Qwen team has released Qwen3.8 Flash, a multimodal reasoning model with a 1 million token context window, according to a listing on OpenRouter. The model is positioned for coding assistance, agentic workflows, visual understanding, document and codebase analysis, desktop interaction, chart analysis, and long-video analysis.

Pricing and Access

Qwen3.8 Flash is available through Alibaba Cloud International at $0.16 per 1M input tokens and $0.47 per 1M output tokens. Caching is supported: cache reads cost $0.016 per 1M tokens, while 5-minute cache writes cost $0.20 per 1M tokens and 5-minute cache reads cost $0.016 per 1M tokens. OpenRouter lists the model's context window at 1 million tokens, positioning it among the longest-context models currently on the platform.

The OpenRouter listing shows a release date of August 26, 2026. No parameter count has been disclosed by Alibaba.

Capabilities

According to Alibaba, Qwen3.8 Flash is a multimodal reasoning model — it accepts multiple input modalities and applies chain-of-thought-style reasoning to tasks. The company positions it for a broad range of use cases: writing and debugging code, running multi-step agentic workflows, interpreting images and charts, parsing long documents and codebases, interacting with desktop environments, and analyzing long-form video content.

No independent benchmark scores are currently available for Qwen3.8 Flash. Alibaba has not published MMLU, HumanEval, or other standard benchmark results alongside this release, and none appear in the OpenRouter listing. Claims about the model's capabilities across coding, agentic tasks, and video analysis remain unverified by third-party testing at this time.

Provider Details

The model is currently served by a single provider — Alibaba Cloud International — with uptime data still being collected. OpenRouter notes latency and throughput data are not yet available, and the listing has been live for approximately three days at time of writing. Requests can be automatically rerouted to alternative healthy providers if an upstream provider errors, subject to a user's routing filters.

What This Means

Qwen3.8 Flash extends Alibaba's Qwen lineup into the "Flash" tier — typically a signal of a smaller, cheaper, faster model meant to compete on cost and latency rather than raw capability ceiling. The $0.16/$0.47 per-million-token pricing undercuts many flagship multimodal models from Western labs, continuing the pattern of Chinese AI labs pricing aggressively to gain API market share. The 1M-token context window, if it holds up under real-world load, would make the model attractive for long-document and long-video workloads where cost per token matters more than peak reasoning quality. The absence of published benchmarks and the lack of provider uptime history mean buyers should treat performance claims as unverified until independent evaluations appear.

Related Articles

model release

Qwen3.8-Flash-Next Debuts with 125B-Parameter Hybrid Architecture, Previews Qwen4 Design

Qwen3.8-Flash-Next is an experimental preview of the architecture Alibaba's Qwen team plans to use for Qwen4, combining hybrid attention, gated residuals, and n-gram embeddings in a 125B-parameter model with only 6B activated per token. Unsloth has released Dynamic 3.0 GGUF quantizations for local inference.

model release

Alibaba Releases Qwen3.8-Flash-Next, a 125B-Parameter Preview of Qwen4's Architecture

Alibaba's Qwen team has released Qwen3.8-Flash-Next, an open-weight model with 125B total parameters (6B activated) that previews architectural changes planned for Qwen4, including a new sparse attention mechanism and n-gram embeddings. The model natively supports 262,144 tokens of context, extensible to 1 million.

model release

Anonymous 'Ox Alpha' Reasoning Model Appears on OpenRouter with Free 1M-Token Context

A stealth model called Ox Alpha has appeared on OpenRouter, offering a 1 million token context window at no cost during its preview period. The model's developer remains anonymous, and OpenRouter says it is acting only as a router, not the model's owner or provider.

model release

Z.ai Releases GLM-5.3 with 1M-Token Context and Always-On Reasoning

Z.ai has released GLM-5.3, a large-scale reasoning model aimed at software engineering and long-horizon agent tasks, featuring a 1M-token context window and mandatory reasoning that cannot be disabled. The model is priced at $1.40 per 1M input tokens and $4.40 per 1M output tokens on OpenRouter.

Comments

Loading...