Qwen3.8 Max Matches Claude Opus 4.8 on Intelligence Index, But Costs 2x More Per Task Than Predecessor
Alibaba's Qwen3.8 Max jumps 10 points to 56 on the Artificial Analysis Intelligence Index, putting it on par with Claude Opus 4.8. But Kimi K3 still edges it out at a lower per-task cost, and Qwen3.8 Max shows a sharp rise in hallucination rate.
Alibaba's Qwen3.8 Max scores 56 on the Artificial Analysis Intelligence Index, according to Artificial Analysis, a 10-point jump over its predecessor Qwen3.7 Max (46). That score puts it on par with Anthropic's Claude Opus 4.8 and ahead of Zhipu AI's GLM-5.2 (51), but it still trails Moonshot AI's Kimi K3, which scores 57 — one point higher — while running 25 percent cheaper per task.
Benchmark gains come with a cost
On GDPval-AA, a benchmark for work-related tasks measured in Elo points, Qwen3.8 Max jumped 468 points to 1,739, moving past Kimi K3 (1,685). Only Claude Opus 5 scores higher, at 1,852.
The gains come with a tradeoff in efficiency. Qwen3.8 Max required 64 steps per task on GDPval-AA, compared to 14 for the previous version. Input token consumption grew 15x, largely because the benchmark resends the full conversation history at each step — a design detail that penalizes models needing more steps to complete a task.
Alibaba cut per-token pricing for Qwen3.8 Max: input tokens dropped from $2.50 to $2.00 per million, output tokens from $7.50 to $6.00, and cached input tokens from $0.50 to $0.25 per million. Despite the lower rates, the increased step count and token volume pushed the effective cost per task on the Intelligence Index to $1.14 — more than double the $0.53 per task for Qwen3.7 Max.
By comparison, Kimi K3 costs $0.86 per task despite scoring one point higher, and GLM-5.2 costs just $0.57 per task. Alibaba's price-to-performance ratio worsened even as list pricing dropped, because the model's higher token consumption outweighed the per-token savings.
Accuracy didn't improve — confidence did
Artificial Analysis also recorded regressions on two benchmarks. AA-LCR, which tests a model's ability to correctly synthesize information from long documents, dropped 2 points compared to Qwen3.7 Max. AA-Omniscience, which measures whether a model answers correctly or honestly declines when it doesn't know, fell 10 points.
The accuracy rate on AA-Omniscience held steady at around 31 percent, but the hallucination rate rose from 23 percent to 40 percent. Qwen3.8 Max answers more often instead of admitting uncertainty, according to Artificial Analysis, even though its actual accuracy didn't improve.
What this means
Qwen3.8 Max illustrates a pattern showing up across frontier model updates: raw intelligence scores can rise while real-world cost efficiency falls. Alibaba's model now matches Claude Opus 4.8 on a composite benchmark, but the path to that score — 64 steps and 15x more input tokens per task — reveals a model that compensates for reasoning gaps with brute-force iteration rather than sharper judgment.
The rising hallucination rate is the more concerning data point. A model that guesses instead of admitting uncertainty 40 percent of the time, up from 23 percent, is a regression in reliability even if headline scores improve. For teams choosing between Qwen3.8 Max, Kimi K3, and GLM-5.2, the Artificial Analysis data suggests Kimi K3 remains the more cost-efficient choice for now, and buyers evaluating any new model release should weigh per-task cost and hallucination rate alongside composite intelligence scores.
Related Articles
Claude Opus 5 Scores 61 on Intelligence Index, Beats Fable 5 on Cost Across Most Benchmarks
Anthropic's Claude Opus 5 posts a 61 on the Artificial Analysis Intelligence Index, narrowly beating Claude Fable 5 (60) and GPT-5.6 Sol (59) while costing less per task. The model leads in coding and knowledge-work benchmarks but shows a rising hallucination rate of 50 percent.
Kimi K3 Scores 32% on Cyber Exploit Benchmark vs. 76% for Leading U.S. Models, Joint UK-US Study Finds
A joint evaluation by the UK AI Security Institute and U.S. Center for AI Standards and Innovation found Kimi K3 scores 32.2% on the ExploitBench benchmark versus 76.2% for leading U.S. models, though it beats China's GLM-5.2 at 24.4%. The gap may stem from Moonshot AI distilling Claude outputs that exclude advanced offensive cyber content.
Frontier AI Models Score Below 50% on First Enterprise IT Benchmark for Kubernetes Incident Response
Artificial Analysis and IBM Research have released ITBench-AA, the first benchmark evaluating AI models on enterprise Site Reliability Engineering tasks. Claude Opus 4.7 leads at 47%, followed by GPT-5.5 at 46% and Qwen3.7 Max at 42%—all frontier models score below 50% on Kubernetes incident response tasks requiring root-cause diagnosis across complex infrastructure.
OpenAI Says GPT-5.6 Sol Scores 38.3% on ARC-AGI-3, Beating Claude Opus 5 — But Only With a Custom Harness
OpenAI says GPT-5.6 Sol scores 38.3 percent on ARC-AGI-3 using a custom Responses API configuration, edging out Claude Opus 5's 30.2 percent. Under the official benchmark harness, GPT-5.6 Sol's score drops to 7.8 percent, raising questions about what the comparison actually measures.
Comments
Loading...