Claude Opus 5 Scores 30.2% on ARC-AGI-3, Nearly 4x the Previous Record
Claude Opus 5 scored 30.2 percent on the ARC-AGI-3 benchmark, nearly four times the previous record of 7.8 percent set by OpenAI's GPT-5.6 Sol (Max). The ARC Prize team attributes the leap to genuinely stronger reasoning, though an independent test on a separate puzzle benchmark showed far smaller improvements.
Opus 5 quadruples the previous ARC-AGI-3 record
Anthropic's Claude Opus 5 has scored 30.2 percent on ARC-AGI-3, according to the ARC Prize Foundation, nearly four times the previous best of 7.8 percent set by OpenAI's GPT-5.6 Sol (Max). The result also surpasses Anthropic's own "Fable-class" models, which reached roughly 20 percent on the same benchmark.
ARC-AGI-3 tests whether a model can infer the rules of an unfamiliar interactive environment, plan a sequence of actions, and execute them step by step — closer to a video game than a static question set. It is designed specifically to resist memorization, since the tasks are meant to be novel relative to training data. Official scores exclude models that rely on external harness software; only the language model's native performance counts.
During testing, Opus 5 solved five of the benchmark's previously unsolved environments, four of them at or above human-level performance. Six of the 25 public demo environments have now been solved in total. ARC Prize has published full results, replays, and benchmarking code.
On older versions of the benchmark suite, Opus 5 scored 90.4 percent on ARC-AGI-2 and 97.5 percent on ARC-AGI-1 — matching prior top scores, though at somewhat higher compute cost, according to ARC Prize.
Researchers observed new problem-solving behavior
ARC Prize's analysis attributes the gain to "genuinely stronger logical reasoning," which it says enables more autonomous exploration, planning, and execution in unfamiliar settings. During testing, Opus 5 displayed behavior not previously documented in prior models, including translating tasks into algebraic notation and independently formulating reflection equations to solve geometric puzzles.
Anthropic has not disclosed the technical basis for the jump. ARC Prize speculates that targeted data labeling combined with reinforcement learning are plausible contributors — Opus 5 is the first Anthropic model developed after ARC-AGI-3's format became public, which could have allowed training efforts to target the benchmark's specific reasoning and puzzle styles without necessarily training on the exact test tasks.
Independent benchmark shows narrower gains
A separate, private benchmark called Witness, built by researcher Guanghan Ning around similar interactive puzzle mechanics, tells a more mixed story. Opus 5 scored 43.4 on Witness, statistically tying Kimi K3 and Fable 5 — a far smaller improvement over Anthropic's own Opus 4.8 than the jump seen on ARC-AGI-3. On one conventional puzzle, Opus 5 correctly identified hidden rules before acting; on a game built around less familiar mechanics, it actually trailed Opus 4.8.
Ning suggested this pattern is consistent with training focused on genre-specific data, though he noted Witness cannot verify what data Anthropic actually used. Greg Kamradt, one of the researchers behind ARC-AGI-3, countered that a single weaker result doesn't offset the model's broader gains, and that Witness's similarity to ARC-AGI-3's format could mean the transfer is real rather than memorized.
Ning later clarified that Opus 5 did generalize to Witness, just to a smaller degree than on ARC-AGI-3, drawing a comparison to how coding benchmarks evolved — from saturated tests like HumanEval toward continuously updated competitions and agentic coding evaluations.
What this means
A 4x jump on a benchmark specifically engineered to resist gaming is a genuine data point, but the discrepancy with Witness scores is a reminder that benchmark-specific optimization and general reasoning improvement can look identical on a single leaderboard. ARC-AGI-3 was public before Opus 5's training, meaning Anthropic had the opportunity — intentionally or not — to shape training data toward the exact skills the benchmark rewards. That doesn't invalidate the result, but it does mean the AI community should watch how Opus 5 performs on the next generation of novel-reasoning benchmarks before treating this as a settled measure of general intelligence. The ARC-AGI-2 and ARC-AGI-1 scores, where Opus 5 merely matches rather than exceeds prior records, suggest the gains may be concentrated specifically in the newer, harder benchmark format.
Related Articles
OpenAI Says GPT-5.6 Sol Scores 38.3% on ARC-AGI-3, Beating Claude Opus 5 — But Only With a Custom Harness
OpenAI says GPT-5.6 Sol scores 38.3 percent on ARC-AGI-3 using a custom Responses API configuration, edging out Claude Opus 5's 30.2 percent. Under the official benchmark harness, GPT-5.6 Sol's score drops to 7.8 percent, raising questions about what the comparison actually measures.
Composio Benchmark: Claude Code Fastest Agent Framework, But Costs Nearly 3x More Than OpenCode
Composio benchmarked DeepSeek V4 Flash across four agent frameworks—Claude Code, Codex, OpenCode, and Oh My Pi—on 30 real-world tasks. Claude Code finished fastest at 122 seconds per task but cost $0.195, nearly three times OpenCode's $0.073, while Oh My Pi had the highest success rate at 17/30 but took 272 seconds per task.
Anthropic's Claude Opus 5 Generates Full 3D Games From a Single Text Prompt, No Assets Required
Anthropic's Claude Opus 5 can generate playable 3D games, including first-person shooters and Minecraft clones, from a single text prompt with zero external assets. Community tests claim it outperforms GPT-5.6 Sol and Kimi K3 in physics realism and mechanical complexity, though no standardized benchmark has confirmed the comparisons.
Anthropic's Claude Opus 4.7 Completes Robot Tasks 20x Faster Than Prior Model, New Benchmark Shows Week-Long Coding Feat
A new Epoch/METR benchmark called MirrorCode shows Claude Opus 4.7 reimplementing large software programs from scratch in tasks estimated to take humans 2-17 weeks, for $251 in inference cost. Separately, Anthropic reports Opus 4.7 completed a suite of quadruped robot tasks in 9 minutes 35 seconds, down from 181 minutes with an earlier model assisting humans.
Comments
Loading...