benchmarkOpenAI

OpenAI's GPT-6 Astra Splits Benchmarks but Beats Human Efficiency on ARC-AGI-3, Pushing Chollet's AGI Timeline Forward

TL;DR

OpenAI's GPT-6 Astra rates first place on Epoch AI's aggregate benchmark but ties its predecessor on Artificial Analysis. Its human-beating move efficiency on ARC-AGI-3 led ARC Prize co-founder François Chollet to call progress 2x faster than expected.

3 min read
0

OpenAI's GPT-6 Astra is producing contradictory results across independent benchmark aggregators, but a single result — human-beating efficiency on ARC-AGI-3 — has led ARC Prize co-founder François Chollet to move up his AGI forecast, which he says is now tracking roughly "2x faster" than he expected.

Two aggregators, two verdicts

Epoch AI combines more than 50 individual benchmarks into its ECI score and ranks GPT-6 Astra first among 267 models tested, at 169 points versus 162 for predecessor GPT-5.6 Sol, 163 for Claude Fable 5.1, and 162 for Claude Opus 5.

Artificial Analysis, which weights knowledge, coding, and text comprehension differently, rates Astra at 61 points — identical to Sol and behind Fable 5.1's 66 and Opus 5's 63.

Pricing tells a similarly split story. OpenAI charges roughly 2.5x more per unit of processed text for Astra than for Sol, making an average task about 75% more expensive. Against Anthropic's Claude Fable 5, however, Astra matches the same coding score at less than half the per-task cost, using roughly a third of the compute steps Sol requires and a fifth of what Opus 5 needs, according to Artificial Analysis.

The ARC-AGI-3 jump

The standout result is on ARC-AGI-3, which drops models into unfamiliar game worlds with no explanation of rules or goals. On ARC Prize's internal harness, Astra scores 62.7%, compared to 7.8% for Sol and 30.2% for Opus 5, at a test cost of roughly $26,000. Under OpenAI's own harness — which preserves reasoning chains across requests and auto-summarizes long runs — the score climbed to a reported 99.9%, running 3.66x faster and using 49% fewer tokens across 167 comparable game-reasoning pairs. ARC Prize says only the internal-harness figure of 62.7% is comparable across vendors.

More notable than the raw score is efficiency: before launch, ARC Prize recorded the median number of moves needed by roughly 500 unscreened human testers to clear each level. Astra cleared 96% of levels in fewer moves than that human median — on average using a little over half as many. Chollet described the behavior on X as "highly efficient, on-the-fly symbolic world modeling for each game and level," noting Astra invents its own algebra-like shorthand (e.g., "extend8 to3; retract10 to2") to track objects, coordinates, and plans — a capability previously seen only in engineered harnesses, not baked into the model itself.

Oddly, reasoning effort and cost move inversely on this benchmark: cost drops from $49,791 at no reasoning to $26,098 at maximum reasoning as the score rises from 35.2% to 62.7%, because Astra solves games in fewer moves and thus fewer model calls. The "low" reasoning setting inexplicably scores 17.5%, worse than no reasoning at all — an anomaly ARC Prize has not explained.

Other results

On ARC-AGI-2, Astra scores 95.0% versus 92.5% (Sol) and 90.4% (Opus 5). On the now-saturated ARC-AGI-1, Astra hits 98.5% at high reasoning effort. On Epoch's FrontierMath Erdős set, Astra became the only model to produce two Lean-verified proofs among 68 open problems at a $300-per-attempt budget (three additional non-standardized solutions cost over $220,000 and don't count toward the score). Hallucination rate on AA-Omniscience dropped from 92% to 51%, but Astra lost roughly 80 Elo points on GDPval-AA v2 and slipped on banking support, SciCode, and long-context tasks. On the Coding Agent Index, Astra scores 67 versus Fable 5.1's 70, using about a third of Sol's token budget.

What this means

The disagreement between Epoch AI and Artificial Analysis underscores how benchmark aggregation methodology, not just raw model capability, now drives headline rankings — Astra's ranking swings from first to tied depending on which tests get weighted. But the ARC-AGI-3 efficiency data is harder to dismiss: it shows a model that not only solves unfamiliar tasks but does so with move-efficiency approaching human performance, using self-invented symbolic notation rather than an external scaffold. That's the specific signal Chollet cites in revising his AGI timeline — not raw accuracy, but the emergence of general-purpose reasoning strategies inside the model weights themselves.

Related Articles

model release

OpenAI's GPT-6 Astra Reportedly Automates AI Engineering Tasks at Under $6 an Hour, According to Latent Space Testing

A Latent Space report describes GPT-6 Astra, a new OpenAI model the blog says can autonomously handle AI engineering tasks—training models, labeling data, deploying systems—at an estimated cost of under $6 per hour. The claims, including 97.6% on FrontierMath and 99.9% on ARC-AGI-3, come from independent blog testing rather than an official OpenAI announcement.

model release

OpenAI Launches GPT-6 Astra, Claims State-of-the-Art Computer Use and 98% on FrontierMath Tier 4

OpenAI has launched GPT-6 Astra, claiming state-of-the-art results on computer use, coding, and scientific reasoning benchmarks. The model is rolling out to a limited set of organizations first, with general ChatGPT availability expected within days.

model release

OpenAI Ships GPT-6 Astra, But Executives Admit They Can't Fully Monitor What It's Thinking

OpenAI released GPT-6 Astra on Thursday, a model president Greg Brockman says could mark the start of AGI. But the model writes out its reasoning less often than prior versions, and OpenAI's chief scientist says monitoring AI thought processes will keep getting harder.

model release

OpenAI Launches GPT-6 Astra, Claims SOTA Computer Use and Coding — But Independent Tests Show Mixed Gains at Higher Cost

OpenAI released GPT-6 Astra on September 3, 2026, claiming state-of-the-art computer use and coding performance alongside new alignment techniques. Independent evaluators found real but uneven gains, higher per-task costs, and reduced chain-of-thought monitorability.

Comments

Loading...