product update

Qwen3.8 Max Prime Launches as Premium High-Throughput SKU at $4/$12 per Million Tokens

TL;DR

Alibaba's Qwen team has launched Qwen3.8 Max Prime, a separate high-throughput SKU of its flagship Qwen3.8 Max model. It carries the same 1M-token context window and multimodal capabilities as the base model but at double the price.

2 min read
0

Alibaba's Qwen team has released Qwen3.8 Max Prime, a higher-throughput variant of its flagship Qwen3.8 Max model, served as a distinct SKU through OpenRouter at a premium price point. The model became available on September 23, 2026.

What's New

Qwen3.8 Max Prime accepts text, image, and video input and returns text output, with a 1-million-token context window and reasoning enabled by default. It supports tool calling, structured outputs, and configurable reasoning effort — the same feature set as the standard Qwen3.8 Max release.

According to Alibaba, the distinguishing factor is throughput: Prime is positioned as a faster-serving option for workloads that need lower latency or higher request volume, rather than a change in underlying capability. Benchmark scores for Prime specifically have not been disclosed.

Pricing

Qwen3.8 Max Prime is priced at $4 per million input tokens and $12 per million output tokens, with cache reads at $0.50 per million tokens through Alibaba Cloud International. That is double the $2/$6 per-million-token pricing of the standard Qwen3.8 Max snapshot (0902), which shares the same 1M context window and multimodal input support.

On OpenRouter, the sole listed provider — Alibaba Cloud International — reports a median (P50) latency of 1.03 seconds and throughput of 51 tokens per second, with 100% uptime and availability over the trailing three-day window.

Context: The Qwen3.8 Max Family

Qwen3.8 Max Prime sits within a broader lineup of Qwen3.8-series releases from Alibaba in 2026. The base Qwen3.8 Max (0902) is a 2.4-trillion-parameter mixture-of-experts model with 95 billion active parameters, post-trained for coding, agentic workflows, and long-document multimodal understanding. An open-weight version, Qwen3.8 2.4T A95B, exposes the same MoE architecture at $2/$6 per million tokens. The family also includes Qwen3.8 Omni Flash for audio-video understanding, Qwen3.8 Flash for lighter multimodal workloads, and Qwen3.8 27B as an open-weight dense vision-language model.

What This Means

Qwen3.8 Max Prime is not a new trained model — it is the same capability set as Qwen3.8 Max repackaged as a premium-tier SKU for customers who need guaranteed higher throughput or lower latency. This mirrors a pattern already common among US labs, where flagship models are offered in multiple serving tiers at different price points depending on infrastructure guarantees rather than differences in model weights.

For developers, the calculus is straightforward: pay double per token for measurably faster, more consistent serving, or stick with the standard Qwen3.8 Max SKU at half the cost if latency is not the binding constraint. The move also signals that Alibaba is treating inference infrastructure as a separate monetization lever from model capability — a strategy that could become more common as context windows and compute costs scale together.

Related Articles

product update

OpenAI Lists GPT-6 Luna Pro: A High-Reasoning Mode for Its Budget GPT-6 Model, Not a New Checkpoint

GPT-6 Luna Pro, listed on OpenRouter with a Sep 22, 2026 release date, is not a distinct model but GPT-6 Luna run with reasoning.mode set to 'pro' for higher-quality outputs on complex tasks. It carries a 1.1M token context window and costs $0.10 per 1M input tokens and $0.50 per 1M output tokens under standard routing.

product update

Amazon Blocks Meta's Muse AI Agent From Making Purchases, Citing Terms of Service

Amazon has blocked Meta's new Muse AI personal agent from completing purchases on its platform, saying the tool violates its terms of service. The move comes as Muse tops Apple's App Store and Meta stock rallies more than 20% in two weeks, with Wall Street watching Zuckerberg's Meta Connect keynote for signs of a broader agentic AI platform strategy.

product update

GitHub Rebuilds Diff Rendering Engine to Open Million-Line Pull Requests in Copilot App

GitHub has re-engineered the diff-viewing surface inside the GitHub Copilot app to support pull requests spanning up to a million lines of code with hundreds of inline comments. The change addresses performance bottlenecks that previously made large-scale code review sluggish or unusable.

product update

AWS Adds Open Weight Models to Amazon Bedrock for Terminal-Based Coding Agents via OpenCode

Amazon Bedrock now supports the open source coding agent OpenCode paired with open weight models including Moonshot AI's Kimi K3, OpenAI's GPT-OSS 120B, and NVIDIA's Nemotron 3 Super 120B. The setup keeps inference inside a customer's AWS account with per-token pricing instead of per-seat subscriptions.

Comments

Loading...