benchmarkAnthropic

Claude Opus 5 Scores 30.2% on ARC-AGI-3, Nearly 4x the Previous Record

TL;DR

Claude Opus 5 scored 30.2 percent on the ARC-AGI-3 benchmark, nearly four times the previous record of 7.8 percent set by OpenAI's GPT-5.6 Sol (Max). The ARC Prize team attributes the leap to genuinely stronger reasoning, though an independent test on a separate puzzle benchmark showed far smaller improvements.

3 min read
0

Opus 5 quadruples the previous ARC-AGI-3 record

Anthropic's Claude Opus 5 has scored 30.2 percent on ARC-AGI-3, according to the ARC Prize Foundation, nearly four times the previous best of 7.8 percent set by OpenAI's GPT-5.6 Sol (Max). The result also surpasses Anthropic's own "Fable-class" models, which reached roughly 20 percent on the same benchmark.

ARC-AGI-3 tests whether a model can infer the rules of an unfamiliar interactive environment, plan a sequence of actions, and execute them step by step — closer to a video game than a static question set. It is designed specifically to resist memorization, since the tasks are meant to be novel relative to training data. Official scores exclude models that rely on external harness software; only the language model's native performance counts.

During testing, Opus 5 solved five of the benchmark's previously unsolved environments, four of them at or above human-level performance. Six of the 25 public demo environments have now been solved in total. ARC Prize has published full results, replays, and benchmarking code.

On older versions of the benchmark suite, Opus 5 scored 90.4 percent on ARC-AGI-2 and 97.5 percent on ARC-AGI-1 — matching prior top scores, though at somewhat higher compute cost, according to ARC Prize.

Researchers observed new problem-solving behavior

ARC Prize's analysis attributes the gain to "genuinely stronger logical reasoning," which it says enables more autonomous exploration, planning, and execution in unfamiliar settings. During testing, Opus 5 displayed behavior not previously documented in prior models, including translating tasks into algebraic notation and independently formulating reflection equations to solve geometric puzzles.

Anthropic has not disclosed the technical basis for the jump. ARC Prize speculates that targeted data labeling combined with reinforcement learning are plausible contributors — Opus 5 is the first Anthropic model developed after ARC-AGI-3's format became public, which could have allowed training efforts to target the benchmark's specific reasoning and puzzle styles without necessarily training on the exact test tasks.

Independent benchmark shows narrower gains

A separate, private benchmark called Witness, built by researcher Guanghan Ning around similar interactive puzzle mechanics, tells a more mixed story. Opus 5 scored 43.4 on Witness, statistically tying Kimi K3 and Fable 5 — a far smaller improvement over Anthropic's own Opus 4.8 than the jump seen on ARC-AGI-3. On one conventional puzzle, Opus 5 correctly identified hidden rules before acting; on a game built around less familiar mechanics, it actually trailed Opus 4.8.

Ning suggested this pattern is consistent with training focused on genre-specific data, though he noted Witness cannot verify what data Anthropic actually used. Greg Kamradt, one of the researchers behind ARC-AGI-3, countered that a single weaker result doesn't offset the model's broader gains, and that Witness's similarity to ARC-AGI-3's format could mean the transfer is real rather than memorized.

Ning later clarified that Opus 5 did generalize to Witness, just to a smaller degree than on ARC-AGI-3, drawing a comparison to how coding benchmarks evolved — from saturated tests like HumanEval toward continuously updated competitions and agentic coding evaluations.

What this means

A 4x jump on a benchmark specifically engineered to resist gaming is a genuine data point, but the discrepancy with Witness scores is a reminder that benchmark-specific optimization and general reasoning improvement can look identical on a single leaderboard. ARC-AGI-3 was public before Opus 5's training, meaning Anthropic had the opportunity — intentionally or not — to shape training data toward the exact skills the benchmark rewards. That doesn't invalidate the result, but it does mean the AI community should watch how Opus 5 performs on the next generation of novel-reasoning benchmarks before treating this as a settled measure of general intelligence. The ARC-AGI-2 and ARC-AGI-1 scores, where Opus 5 merely matches rather than exceeds prior records, suggest the gains may be concentrated specifically in the newer, harder benchmark format.

Related Articles

benchmark

OpenAI's GPT-6 Astra Splits Benchmarks but Beats Human Efficiency on ARC-AGI-3, Pushing Chollet's AGI Timeline Forward

OpenAI's GPT-6 Astra rates first place on Epoch AI's aggregate benchmark but ties its predecessor on Artificial Analysis. Its human-beating move efficiency on ARC-AGI-3 led ARC Prize co-founder François Chollet to call progress 2x faster than expected.

research

Anthropic's Claude Fable 5.1 Reportedly Solves 1653 Royalist Cipher in 44 Minutes

According to testing firm Vals AI, Anthropic's Claude Fable 5.1 independently identified and solved the 'Cyphral Distich,' a 1653 numeric cipher by Sir Thomas Urquhart that had defeated other frontier models. The AI decoded a hidden pro-royalist message by mapping each number to a word in Urquhart's original text.

research

Anthropic Joins Google in Watermarking AI-Generated Text, Reviving Debate Over Output Quality

Anthropic announced on August 11 that all future Claude models will embed an invisible watermark in generated text, following Google's lead with SynthID-Text. The move is partly driven by the EU AI Act, which mandates watermarking for AI models released after August 2, 2026, though researchers remain split on whether the technique degrades output quality.

benchmark

Artificial Analysis Updates Intelligence Index to v4.2, Narrows GPT-6 Astra Gap Controversy

Artificial Analysis released version 4.2 of its Intelligence Index after its original scoring showed GPT-6 Astra barely improving on its predecessor, contradicting Epoch AI's ranking of Astra as the top model out of 267 tested. The update adds two benchmarks, drops the saturated GPQA-Diamond, and increases private test weighting to 40 percent.

Comments

Loading...

Claude Opus 5 Sets New ARC-AGI-3 Record at 30.2% | TPS