Qwen3.8-2.4T-A95B

Alibaba / Qwen🇨🇳 China
active
Context window262K tokens
Input / 1M tokens$2
Output / 1M tokens$6

Version History

Qwen3.8major

Alibaba released Qwen3.8-2.4T-A95B-FP8, an open-weight, FP8-quantized MoE checkpoint with 2.4T total and 95B activated parameters, serving as the foundation for the hosted Qwen3.8-Max API. The model extends context to 1M tokens and shows significant benchmark gains over Qwen3.7-Max on coding and agentic tasks, according to Alibaba.

3.8major

Alibaba released Qwen3.8-2.4T-A95B, a 2.4-trillion-parameter MoE model with 95B active parameters, built on the Qwen3.5 architecture with hybrid Gated DeltaNet and Gated Attention layers. It introduces reasoning_effort and preserve_thinking controls and claims Qwen-Max-class performance in an open-weight release for the first time.

Benchmark Scores

Full leaderboard →
92.6%
GPQA
37.0 tokens_per_sec
Speed (tok/s)

Coverage