OpenAI Says GPT-5.6 Sol Scores 38.3% on ARC-AGI-3, Beating Claude Opus 5 — But Only With a Custom Harness
OpenAI says GPT-5.6 Sol scores 38.3 percent on ARC-AGI-3 using a custom Responses API configuration, edging out Claude Opus 5's 30.2 percent. Under the official benchmark harness, GPT-5.6 Sol's score drops to 7.8 percent, raising questions about what the comparison actually measures.
What happened
OpenAI claims GPT-5.6 Sol scores 38.3 percent on the ARC-AGI-3 benchmark, surpassing Claude Opus 5's 30.2 percent score. The catch: OpenAI's number comes from a custom test setup, not the official ARC-AGI-3 harness that Anthropic used to post its result.
Anthropic's Opus 5 quadrupled the previous ARC-AGI-3 record when it launched, hitting 30.2 percent under the benchmark's standard constraints. ARC-AGI-3 is designed to test raw model reasoning without external scaffolding or tooling assistance.
OpenAI ran GPT-5.6 Sol through its own Responses API using two non-standard features:
- Retained Reasoning: keeps the model's chain of thought active between steps instead of discarding it after each action
- Compaction: summarizes prior context rather than truncating it when the context window fills up
With both features enabled, GPT-5.6 Sol reached 38.3 percent, according to OpenAI. Under the official ARC-AGI-3 test harness — which strips reasoning traces after each action — the same model scored just 7.8 percent.
The core issue
OpenAI argues that benchmark scores reflect not just a model's raw capability but the technical infrastructure surrounding it, and that judging models only on a stripped-down harness undersells what they can do in production settings. That argument has merit in general, but it cuts against the specific design goal of ARC-AGI-3, which intentionally removes external scaffolding to isolate pure reasoning performance rather than measure how a model performs inside a favorable API configuration.
Anthropic's 30.2 percent score for Opus 5 was produced under those same standard constraints — no retained reasoning, no custom compaction. If Opus 5 were run inside Claude Code or a similarly optimized harness, its score would likely rise as well, though Anthropic has not published such a number.
No independent lab has verified either the 38.3 percent or 7.8 percent figures for GPT-5.6 Sol. Both numbers come directly from OpenAI, and Opus 5's 30.2 percent figure comes directly from Anthropic. Neither company's benchmark claims have been reproduced by a third party as of this writing.
What this means
The headline comparison — 38.3 percent versus 30.2 percent — is not measuring the same thing. OpenAI's number reflects GPT-5.6 Sol plus a custom reasoning-retention and context-compaction pipeline; Anthropic's number reflects Opus 5 running under the benchmark's default, no-scaffolding conditions. Under those same default conditions, GPT-5.6 Sol scored 7.8 percent — less than a third of Opus 5's result.
This pattern is becoming common in frontier-model benchmarking: vendors increasingly report scores from tuned harnesses rather than standardized ones, making cross-model comparisons unreliable without checking the fine print. For anyone evaluating these models for real deployments, the more useful signal here may be the gap itself — 7.8 percent to 38.3 percent — which shows how heavily ARC-AGI-3 performance depends on infrastructure choices like reasoning retention and context management, independent of the underlying model weights.
Related Articles
AWS Benchmark: OpenAI's GPT-5.6 Luna Beats GPT-5.4 Mini on Cost-Per-Correct-Answer Despite Similar List Price
An AWS blog post using an open-source benchmarking harness finds that GPT-5.6 Luna, Terra, and Sol on Amazon Bedrock deliver lower cost-per-correct-answer than OpenAI's cost-optimized GPT-5.4 Mini and Nano, once accuracy, token efficiency, and agent turn counts are factored in. The analysis also cites a July 30, 2026 price cut of up to 80% for GPT-5.6 Luna on Amazon Bedrock.
OpenAI's GPT-6 Astra Beats Claude Fable 5.1 Nearly 3-to-1 in Autonomous Business Benchmark, Tops Drone Navigation Tests
Independent testing lab Andon Labs found OpenAI's GPT-6 Astra nearly triples Claude Fable 5.1's performance running a simulated vending machine business, averaging $15,515 versus $5,422. Astra also became the first model to beat human-AI baseline performance across all five Drone-Bench subtasks, including autonomous person-tracking via drone.
GPT-6 Astra Beats Ai2's MolmoAct2 on New Robotics Benchmark, Researcher Calls It a 'Step Change'
A new robotics benchmark called StationeryBench shows OpenAI's GPT-6 Astra completing 7 of 100 desk-object manipulation tasks versus zero for Ai2's MolmoAct2, with a median progress score of 46 against 12. Cornell/DeepMind researcher Yoav Artzi calls the result a 'step change in spatial reasoning.'
OpenAI Claims 10,000-Agent System Solved Navier-Stokes Problem in 88 Hours; Mathematician Disputes Independence of Resul
OpenAI claims a system of roughly 10,000 coordinating AI agents produced a solution to the Navier-Stokes equations, one of seven unsolved Millennium Prize Problems, in 88 hours. NYU mathematician Tristan Buckmaster has publicly questioned whether OpenAI's approach drew on his own unpublished work with Anthropic researcher Levent Alpöge.
Comments
Loading...