benchmark

Artificial Analysis Updates Intelligence Index to v4.2, Narrows GPT-6 Astra Gap Controversy

TL;DR

Artificial Analysis released version 4.2 of its Intelligence Index after its original scoring showed GPT-6 Astra barely improving on its predecessor, contradicting Epoch AI's ranking of Astra as the top model out of 267 tested. The update adds two benchmarks, drops the saturated GPQA-Diamond, and increases private test weighting to 40 percent.

3 min read
0

Artificial Analysis has released version 4.2 of its Intelligence Index following criticism that its previous scoring failed to capture real-world gains in OpenAI's GPT-6 Astra model.

The original index scored Astra as roughly on par with its predecessor, Sol, despite other evaluations showing much larger jumps. Epoch AI ranked Astra first out of 267 models tested, with 169 points across more than 50 benchmarks. ARC-AGI-3 results also showed a large to very large improvement over Sol, depending on which harness was used to run the test. OpenAI's own internal evaluations placed Astra well ahead of competing models.

What changed in v4.2

Under the revised index, GPT-6 Astra now shows a four-point gain over Sol — the two models were previously tied. Claude Fable 5.1 from Anthropic still leads the overall ranking, with Astra in second place and a Meta model in third.

Artificial Analysis also reports that Astra uses fewer tokens per task than any other frontier model it has tested, a claim the company presents as a separate efficiency metric alongside raw capability scores. On cost-to-performance ratio, Anthropic, OpenAI, Meta, and Zhipu AI now share the top position, according to Artificial Analysis's figures.

The update makes several methodology changes beyond re-scoring Astra:

  • Two new benchmarks were added: AA-Briefcase, which tests real-world knowledge work, and GDP.pdf from Surge AI, which evaluates PDF document analysis.
  • GPQA-Diamond was dropped from the index because, according to Artificial Analysis, frontier models have effectively solved it, making it no longer useful for differentiating performance.
  • Private test data — questions not publicly available for training — now accounts for 40 percent of the overall weighting, up from a smaller share previously, intended to make it harder for labs to optimize models specifically for the benchmark.
  • Artificial Analysis says it also fixed unspecified scoring errors across several existing benchmarks and adjusted its grading systems to produce more stable results.

Why the update happened now

Artificial Analysis says it typically avoids changing its methodology around major model launches to keep scores comparable over time. But the company says the pace of top-of-leaderboard movement made an interim update necessary before its next full revision. That larger overhaul, referred to as version 5, has reportedly been in development for eight months and will roll out in stages rather than as a single release.

What this means

The episode is a reminder that third-party benchmark aggregators are themselves subject to methodology risk — a scoring system that doesn't update its test mix or weighting can drift out of sync with what models are actually capable of, especially as new frontier releases target capabilities the original benchmark wasn't designed to measure. The gap between Epoch AI's ranking (Astra first out of 267 models) and Artificial Analysis's original near-tie is a concrete illustration of how much a benchmark's specific test selection and harness choice can swing relative rankings between labs. Increasing private test weighting to 40 percent is a direct response to benchmark gaming concerns, where labs may fine-tune on or near public eval questions. For engineers choosing between models based on published leaderboards, this is a caution against treating any single index as ground truth — cross-referencing multiple evaluation sources, as happened here, is what surfaced the discrepancy in the first place.

Related Articles

benchmark

Claude Opus 5 Scores 61 on Intelligence Index, Beats Fable 5 on Cost Across Most Benchmarks

Anthropic's Claude Opus 5 posts a 61 on the Artificial Analysis Intelligence Index, narrowly beating Claude Fable 5 (60) and GPT-5.6 Sol (59) while costing less per task. The model leads in coding and knowledge-work benchmarks but shows a rising hallucination rate of 50 percent.

benchmark

Moonshot AI's Kimi K3 matches top US models at 40% lower cost, will be open-weight

Moonshot AI's Kimi K3 model has matched or exceeded performance of Anthropic's Opus 4.8 and OpenAI's GPT-5.6 Sol in independent benchmarks while costing 40% less than comparable US models. The Beijing-based company plans to release Kimi K3 as an open-weight model on July 27.

benchmark

GLM-5.3 Ties Kimi K3 for Top Open-Model Ranking, Undercuts Rivals on Price — But Open Weights Delayed

Z.ai's GLM-5.3 ties Kimi K3 for the top spot among open models on the Artificial Analysis Intelligence Index, driven by a major leap in agentic task performance. The company is delaying the open-weight release by about two weeks, citing the model's unusually strong vulnerability-detection capabilities.

benchmark

Artificial Analysis Launches Optima, a Platform to Build Custom AI Benchmarks on Your Own Data

Artificial Analysis has launched Optima, a platform that lets users build custom AI benchmarks using their own data, workflows, or use-case descriptions. Unlike public benchmarks, Optima compares models on cost per task and time per task in addition to quality.

Comments

Loading...