Claude Opus 5 Scores 30.2% on ARC-AGI-3, Nearly 4x the Previous Record
Claude Opus 5 scored 30.2 percent on the ARC-AGI-3 benchmark, nearly four times the previous record of 7.8 percent set by OpenAI's GPT-5.6 Sol (Max). The ARC Prize team attributes the leap to genuinely stronger reasoning, though an independent test on a separate puzzle benchmark showed far smaller improvements.
Opus 5 quadruples the previous ARC-AGI-3 record
Anthropic's Claude Opus 5 has scored 30.2 percent on ARC-AGI-3, according to the ARC Prize Foundation, nearly four times the previous best of 7.8 percent set by OpenAI's GPT-5.6 Sol (Max). The result also surpasses Anthropic's own "Fable-class" models, which reached roughly 20 percent on the same benchmark.
ARC-AGI-3 tests whether a model can infer the rules of an unfamiliar interactive environment, plan a sequence of actions, and execute them step by step — closer to a video game than a static question set. It is designed specifically to resist memorization, since the tasks are meant to be novel relative to training data. Official scores exclude models that rely on external harness software; only the language model's native performance counts.
During testing, Opus 5 solved five of the benchmark's previously unsolved environments, four of them at or above human-level performance. Six of the 25 public demo environments have now been solved in total. ARC Prize has published full results, replays, and benchmarking code.
On older versions of the benchmark suite, Opus 5 scored 90.4 percent on ARC-AGI-2 and 97.5 percent on ARC-AGI-1 — matching prior top scores, though at somewhat higher compute cost, according to ARC Prize.
Researchers observed new problem-solving behavior
ARC Prize's analysis attributes the gain to "genuinely stronger logical reasoning," which it says enables more autonomous exploration, planning, and execution in unfamiliar settings. During testing, Opus 5 displayed behavior not previously documented in prior models, including translating tasks into algebraic notation and independently formulating reflection equations to solve geometric puzzles.
Anthropic has not disclosed the technical basis for the jump. ARC Prize speculates that targeted data labeling combined with reinforcement learning are plausible contributors — Opus 5 is the first Anthropic model developed after ARC-AGI-3's format became public, which could have allowed training efforts to target the benchmark's specific reasoning and puzzle styles without necessarily training on the exact test tasks.
Independent benchmark shows narrower gains
A separate, private benchmark called Witness, built by researcher Guanghan Ning around similar interactive puzzle mechanics, tells a more mixed story. Opus 5 scored 43.4 on Witness, statistically tying Kimi K3 and Fable 5 — a far smaller improvement over Anthropic's own Opus 4.8 than the jump seen on ARC-AGI-3. On one conventional puzzle, Opus 5 correctly identified hidden rules before acting; on a game built around less familiar mechanics, it actually trailed Opus 4.8.
Ning suggested this pattern is consistent with training focused on genre-specific data, though he noted Witness cannot verify what data Anthropic actually used. Greg Kamradt, one of the researchers behind ARC-AGI-3, countered that a single weaker result doesn't offset the model's broader gains, and that Witness's similarity to ARC-AGI-3's format could mean the transfer is real rather than memorized.
Ning later clarified that Opus 5 did generalize to Witness, just to a smaller degree than on ARC-AGI-3, drawing a comparison to how coding benchmarks evolved — from saturated tests like HumanEval toward continuously updated competitions and agentic coding evaluations.
What this means
A 4x jump on a benchmark specifically engineered to resist gaming is a genuine data point, but the discrepancy with Witness scores is a reminder that benchmark-specific optimization and general reasoning improvement can look identical on a single leaderboard. ARC-AGI-3 was public before Opus 5's training, meaning Anthropic had the opportunity — intentionally or not — to shape training data toward the exact skills the benchmark rewards. That doesn't invalidate the result, but it does mean the AI community should watch how Opus 5 performs on the next generation of novel-reasoning benchmarks before treating this as a settled measure of general intelligence. The ARC-AGI-2 and ARC-AGI-1 scores, where Opus 5 merely matches rather than exceeds prior records, suggest the gains may be concentrated specifically in the newer, harder benchmark format.
Related Articles
Anthropic's Claude Opus 5 Hits 0% Prompt Injection Success Rate in Browser Agent Tests, With Defenses Enabled
Anthropic's system card for Claude Opus 5 reports a 0% prompt injection success rate across 129 browser agent test scenarios when Auto Mode is enabled. On Gray Swan's broader indirect prompt injection benchmark, Opus 5 posted a 2.0% attacker success rate after 15 attempts, the lowest among tested frontier models.
Claude Opus 5 Scores 61 on Intelligence Index, Beats Fable 5 on Cost Across Most Benchmarks
Anthropic's Claude Opus 5 posts a 61 on the Artificial Analysis Intelligence Index, narrowly beating Claude Fable 5 (60) and GPT-5.6 Sol (59) while costing less per task. The model leads in coding and knowledge-work benchmarks but shows a rising hallucination rate of 50 percent.
Anthropic Ships Claude Opus 5, Claims Near-Fable Performance at Half the Price
Anthropic released Claude Opus 5 on July 24, 2026, positioning it as a lower-cost alternative to its more expensive Claude Fable 5 model. Independent evaluators Epoch AI and Artificial Analysis report mixed but largely favorable results, with Opus 5 nearly matching Fable 5 on coding benchmarks while cutting cost-per-task by roughly 20%.
Anthropic Ships Claude Opus 5, Claims It Matches Flagship Fable 5 on Coding at Half the Cost
Anthropic released Claude Opus 5 on July 24, its fourth model launch in under two months, priced at $5 per million input tokens and $25 per million output tokens. The company claims the model matches or beats its flagship Fable 5 on most coding and knowledge-work benchmarks while posting the lowest deception rate of any model it has shipped.
Comments
Loading...