benchmark

Qwen3.8 Max Matches Claude Opus 4.8 on Intelligence Index, But Costs 2x More Per Task Than Predecessor

TL;DR

Alibaba's Qwen3.8 Max jumps 10 points to 56 on the Artificial Analysis Intelligence Index, putting it on par with Claude Opus 4.8. But Kimi K3 still edges it out at a lower per-task cost, and Qwen3.8 Max shows a sharp rise in hallucination rate.

3 min read
0

Alibaba's Qwen3.8 Max scores 56 on the Artificial Analysis Intelligence Index, according to Artificial Analysis, a 10-point jump over its predecessor Qwen3.7 Max (46). That score puts it on par with Anthropic's Claude Opus 4.8 and ahead of Zhipu AI's GLM-5.2 (51), but it still trails Moonshot AI's Kimi K3, which scores 57 — one point higher — while running 25 percent cheaper per task.

Benchmark gains come with a cost

On GDPval-AA, a benchmark for work-related tasks measured in Elo points, Qwen3.8 Max jumped 468 points to 1,739, moving past Kimi K3 (1,685). Only Claude Opus 5 scores higher, at 1,852.

The gains come with a tradeoff in efficiency. Qwen3.8 Max required 64 steps per task on GDPval-AA, compared to 14 for the previous version. Input token consumption grew 15x, largely because the benchmark resends the full conversation history at each step — a design detail that penalizes models needing more steps to complete a task.

Alibaba cut per-token pricing for Qwen3.8 Max: input tokens dropped from $2.50 to $2.00 per million, output tokens from $7.50 to $6.00, and cached input tokens from $0.50 to $0.25 per million. Despite the lower rates, the increased step count and token volume pushed the effective cost per task on the Intelligence Index to $1.14 — more than double the $0.53 per task for Qwen3.7 Max.

By comparison, Kimi K3 costs $0.86 per task despite scoring one point higher, and GLM-5.2 costs just $0.57 per task. Alibaba's price-to-performance ratio worsened even as list pricing dropped, because the model's higher token consumption outweighed the per-token savings.

Accuracy didn't improve — confidence did

Artificial Analysis also recorded regressions on two benchmarks. AA-LCR, which tests a model's ability to correctly synthesize information from long documents, dropped 2 points compared to Qwen3.7 Max. AA-Omniscience, which measures whether a model answers correctly or honestly declines when it doesn't know, fell 10 points.

The accuracy rate on AA-Omniscience held steady at around 31 percent, but the hallucination rate rose from 23 percent to 40 percent. Qwen3.8 Max answers more often instead of admitting uncertainty, according to Artificial Analysis, even though its actual accuracy didn't improve.

What this means

Qwen3.8 Max illustrates a pattern showing up across frontier model updates: raw intelligence scores can rise while real-world cost efficiency falls. Alibaba's model now matches Claude Opus 4.8 on a composite benchmark, but the path to that score — 64 steps and 15x more input tokens per task — reveals a model that compensates for reasoning gaps with brute-force iteration rather than sharper judgment.

The rising hallucination rate is the more concerning data point. A model that guesses instead of admitting uncertainty 40 percent of the time, up from 23 percent, is a regression in reliability even if headline scores improve. For teams choosing between Qwen3.8 Max, Kimi K3, and GLM-5.2, the Artificial Analysis data suggests Kimi K3 remains the more cost-efficient choice for now, and buyers evaluating any new model release should weigh per-task cost and hallucination rate alongside composite intelligence scores.

Related Articles

benchmark

AWS Benchmark: OpenAI's GPT-5.6 Luna Beats GPT-5.4 Mini on Cost-Per-Correct-Answer Despite Similar List Price

An AWS blog post using an open-source benchmarking harness finds that GPT-5.6 Luna, Terra, and Sol on Amazon Bedrock deliver lower cost-per-correct-answer than OpenAI's cost-optimized GPT-5.4 Mini and Nano, once accuracy, token efficiency, and agent turn counts are factored in. The analysis also cites a July 30, 2026 price cut of up to 80% for GPT-5.6 Luna on Amazon Bedrock.

benchmark

GLM-5.3 Ties Kimi K3 for Top Open-Model Ranking, Undercuts Rivals on Price — But Open Weights Delayed

Z.ai's GLM-5.3 ties Kimi K3 for the top spot among open models on the Artificial Analysis Intelligence Index, driven by a major leap in agentic task performance. The company is delaying the open-weight release by about two weeks, citing the model's unusually strong vulnerability-detection capabilities.

benchmark

Artificial Analysis Launches Search Index Benchmark for AI Agent Search APIs

Artificial Analysis has released the Search Index, a benchmark measuring how search API providers perform for AI agents across quality, cost, and speed. Parallel, Exa, and Firecrawl lead the initial rankings, with search access boosting model scores from 33 to as high as 75 points.

benchmark

Claude Opus 5 Scores 61 on Intelligence Index, Beats Fable 5 on Cost Across Most Benchmarks

Anthropic's Claude Opus 5 posts a 61 on the Artificial Analysis Intelligence Index, narrowly beating Claude Fable 5 (60) and GPT-5.6 Sol (59) while costing less per task. The model leads in coding and knowledge-work benchmarks but shows a rising hallucination rate of 50 percent.

Comments

Loading...