OpenAI's GPT-6 Astra Tops Math Benchmark Despite Deliberately Skipping Math Optimization
OpenAI's GPT-6 Astra took first place on ulam.ai's ErdosBench, solving 106 of 226 open math problems and disproving 27 others. Chief scientist Jakub Pachocki says the company deliberately didn't optimize for math research, prioritizing recursive self-improvement work instead.
OpenAI's GPT-6 Astra has taken the top spot on ErdosBench, a benchmark of 226 open math problems modeled on the famous Erdős problems, scoring 3.23 and solving 106 problems outright — 43 of them completely. The model also disproved 27 additional problems.
That performance beat GPT-5.6 Sol, which scored 3.12 and solved 78 problems at maximum reasoning settings on the same benchmark, according to ulam.ai. Benchmark developer Przemek Chojecki characterized the jump as "a solid 5%-10% gain on various math-research skills tested," while noting the benchmark remains far from saturated.
What makes the result notable isn't just the score — it's that OpenAI says it didn't try to maximize it.
OpenAI chose not to optimize for math
In an essay titled "An Alien Mind," OpenAI chief scientist Jakub Pachocki wrote that the company "could make the models better at specifically mathematics research with additional focus," but chose not to prioritize that direction. Instead, according to Pachocki, OpenAI is concentrating resources on recursive self-improvement (RSI) and automated alignment research, which he describes as "the only way to remain at the frontier of AI research moving forward."
That statement means GPT-6 Astra's math performance is a byproduct of broader capability gains rather than a targeted push. It also confirms something researchers have suspected: even OpenAI faces hard trade-offs at the frontier. Pushing one capability area — in this case, math — means pulling resources from elsewhere.
Compared to Sol, Astra reportedly showed stronger scientific writing quality and was less prone to overstating its findings, in some cases understating its own results, according to the source reporting on the benchmark.
A spiky path to capability, not a broad one
The result feeds into a broader debate about how AI capability is actually developing. Cambridge researcher Adam Hunt has visualized two competing trajectories: a "mainstream AGI" thesis where models improve broadly and gradually across all human tasks, versus a "spiky" trajectory where models achieve extreme strength in narrow domains like coding and math while stagnating or regressing in areas like language quality, common sense, or social reasoning.
Pachocki's admission that Astra's math strength is incidental — not the product of deliberate optimization — is being read as evidence for the spiky thesis. If even the leading lab can't advance every capability simultaneously, general intelligence in the broadest sense remains further off than raw benchmark scores suggest.
Mathematicians face a different problem: too many proofs
At the 2026 International Congress of Mathematicians, mathematician Terence Tao raised a separate concern: as AI models generate proofs faster than humans can verify them, the field risks shifting from a scarcity of solutions to an overload of unverified results. The bottleneck, in Tao's view, could shift from solving problems to deciding which results are actually significant — a shift he compares to the foundational crisis mathematics faced in the early 20th century.
The hardest open problems in mathematics remain unsolved, buying the field time. But skepticism persists among mathematicians who doubt language models can produce genuinely creative breakthroughs rather than incremental extensions of existing techniques.
What this means
GPT-6 Astra's ErdosBench win demonstrates that even unoptimized models from frontier labs now outperform specialized math tools on hard research problems — a sign of broad capability gains rather than narrow tuning. But OpenAI's own account undercuts any narrative of steady, uniform progress toward general intelligence: resources are being explicitly redirected toward recursive self-improvement and alignment work, not mathematics. That's a bet that automating AI research itself will eventually yield gains everywhere, including math, faster than direct optimization would. Whether that bet pays off — and whether verification, not generation, becomes math's real bottleneck — remains an open question the field is only beginning to confront.
Related Articles
OpenAI Claims Resolution to Navier–Stokes Millennium Prize Problem Amid Priority Dispute
OpenAI claims its unreleased internal model resolved the Navier–Stokes existence and smoothness problem, one of seven $1 million Millennium Prize Problems. NYU professor Tristan Buckmaster disputes the timeline, alleging OpenAI moved after hearing rumors of his own near-year-long collaboration with Anthropic's Levent Alpöge.
OpenAI Claims Unreleased Model Solved Navier-Stokes Problem, Faces Credit Dispute
OpenAI claims an internal, unreleased model solved part of the Navier-Stokes Millennium Prize problem using roughly 10,000 concurrent agents over 88 hours. NYU mathematician Tristan Buckmaster has questioned whether the effort drew on his unpublished research with an Anthropic researcher.
Unverified 'GPT-6 Astra' Reportedly Completes Portal Solo in Under 24 Hours, No Official OpenAI Confirmation
A developer named cozyblaze posted on X that a model called 'GPT-6 Astra' completed Portal from start to finish without human intervention in 23 hours 43 minutes. OpenAI has not confirmed the existence of a model by that name, and all details come from a single third-party account.
OpenAI Claims Unreleased Model Solved Navier-Stokes Millennium Prize Problem in 88 Hours, Faces Scooping Allegations
OpenAI announced its unreleased internal model solved the Navier-Stokes Millennium Prize problem in 88 hours using roughly 10,000 AI agents, but the timing—one day after related findings from NYU and Anthropic researchers—has triggered allegations of scooping and improper data access. OpenAI denies using specific user data but cannot rule out indirect influence from de-identified usage data.
Comments
Loading...