model release

Alibaba Releases Qwen3.8 Flash, a Multimodal Reasoning Model with 1M-Token Context

TL;DR

Alibaba has released Qwen3.8 Flash, a multimodal reasoning model with a 1 million token context window, aimed at coding, agentic workflows, and visual/document analysis. It's priced at $0.16 per 1M input tokens and $0.47 per 1M output tokens through Alibaba Cloud International.

2 min read
0

Qwen3.8 Flash — Quick Specs

Context window1000K tokens
Input$0.16/1M tokens
Output$0.47/1M tokens

Alibaba's Qwen team has released Qwen3.8 Flash, a multimodal reasoning model with a 1 million token context window, according to a listing on OpenRouter. The model is positioned for coding assistance, agentic workflows, visual understanding, document and codebase analysis, desktop interaction, chart analysis, and long-video analysis.

Pricing and Access

Qwen3.8 Flash is available through Alibaba Cloud International at $0.16 per 1M input tokens and $0.47 per 1M output tokens. Caching is supported: cache reads cost $0.016 per 1M tokens, while 5-minute cache writes cost $0.20 per 1M tokens and 5-minute cache reads cost $0.016 per 1M tokens. OpenRouter lists the model's context window at 1 million tokens, positioning it among the longest-context models currently on the platform.

The OpenRouter listing shows a release date of August 26, 2026. No parameter count has been disclosed by Alibaba.

Capabilities

According to Alibaba, Qwen3.8 Flash is a multimodal reasoning model — it accepts multiple input modalities and applies chain-of-thought-style reasoning to tasks. The company positions it for a broad range of use cases: writing and debugging code, running multi-step agentic workflows, interpreting images and charts, parsing long documents and codebases, interacting with desktop environments, and analyzing long-form video content.

No independent benchmark scores are currently available for Qwen3.8 Flash. Alibaba has not published MMLU, HumanEval, or other standard benchmark results alongside this release, and none appear in the OpenRouter listing. Claims about the model's capabilities across coding, agentic tasks, and video analysis remain unverified by third-party testing at this time.

Provider Details

The model is currently served by a single provider — Alibaba Cloud International — with uptime data still being collected. OpenRouter notes latency and throughput data are not yet available, and the listing has been live for approximately three days at time of writing. Requests can be automatically rerouted to alternative healthy providers if an upstream provider errors, subject to a user's routing filters.

What This Means

Qwen3.8 Flash extends Alibaba's Qwen lineup into the "Flash" tier — typically a signal of a smaller, cheaper, faster model meant to compete on cost and latency rather than raw capability ceiling. The $0.16/$0.47 per-million-token pricing undercuts many flagship multimodal models from Western labs, continuing the pattern of Chinese AI labs pricing aggressively to gain API market share. The 1M-token context window, if it holds up under real-world load, would make the model attractive for long-document and long-video workloads where cost per token matters more than peak reasoning quality. The absence of published benchmarks and the lack of provider uptime history mean buyers should treat performance claims as unverified until independent evaluations appear.

Related Articles

model release

Qwen releases Qwen-Image-2.1-Turbo: 8-step text-to-image and editing checkpoint on a 7B architecture

Qwen has published Qwen-Image-2.1-Turbo on Hugging Face, an accelerated checkpoint of Qwen-Image-2.1 that runs text-to-image generation and image editing in 8 denoising steps. It keeps the same 7B visual generation architecture and loads through a new QwenImage21Pipeline in Diffusers.

model release

StepFun releases Step 5 Preview: 600B MoE with 1M context at $1/$2.70 per 1M tokens

StepFun has listed Step 5 Preview, a sparse Mixture-of-Experts model with 600B total and 27B active parameters and a 1.0M-token context window. It is priced at $1 input and $2.70 output per 1M tokens on OpenRouter. StepFun positions it as its flagship model for agentic work.

model release

Google releases Nano Banana 2.1 image model: $1.50/$30 per 1M tokens, 66K context

Google's Nano Banana 2.1 (Gemini Nano Banana 2.1) is an image generation and editing model on the Flash tier, listed on OpenRouter at $1.50 input and $30 output per 1M tokens with a 66K context window. It supports 1K, 2K, and 4K output and succeeds Nano Banana 2 and Nano Banana Pro, according to the listing.

model release

Microsoft releases Decision-1, a Qwen3.5-9B-based model for classification and routing, at $0.042 per 1M input tokens

Microsoft has released Decision-1, a decision model built on Qwen3.5-9B for classification, evaluation, and routing. Microsoft claims 83.5% accuracy across 36 benchmarks and 85 ms latency. Input tokens cost $0.042 per million, and output tokens are free.

Comments

Loading...