OpenAI Claims Internal Astra Model Solved 10 Decade-Old Math Problems for Under $2,000 Each
OpenAI claims an internal version of its next major model, Astra, produced solutions to ten mathematical and theoretical computer science problems that had seen no progress in at least a decade. The company says each solution cost less than $2,000 in GPT-5.6 Sol token pricing, and published Lean 4 formalizations along with a paper describing the results.
OpenAI claims an internal version of its upcoming Astra model produced verified solutions to ten mathematical and theoretical computer science problems, each of which had reportedly seen no progress on its main result for at least a decade. According to OpenAI, each solution cost less than $2,000 in token spending at GPT-5.6 Sol pricing.
The company has published Lean 4 formalizations of the results in the openai/ten-proofs GitHub repository, alongside a paper detailing the solutions. OpenAI also released an LLM-generated PDF in which the model, according to OpenAI, "reconstructs how the proof came together" based on reasoning traces that were not themselves published.
No pricing or release date for Astra itself has been disclosed. OpenAI has not stated how many problems it attempted without success, nor how much was spent on unsuccessful attempts — a gap independent observers have noted, since the $2,000-per-problem figure only covers the ten reported wins.
Context: a pattern of AI math research claims
This announcement follows a similar move by Anthropic days earlier, when the company said it used a Mythos Preview system to discover cryptographic weaknesses with Claude, reportedly spending $100,000 on tokens. Anthropic's prompts, according to reporting on the effort, explicitly instructed the model to avoid "low hanging fruit" and pursue "genuinely hard findings."
The OpenAI release includes Lean 4 formal proofs, which allow independent verification through the Lean proof assistant rather than requiring trust in the model's claims alone. This is a meaningfully higher bar for transparency than a plain assertion of results, though the underlying prompts used to generate the solutions have not been released.
Mathematician Terence Tao described a related concept in IEEE Spectrum in June, calling it "big mathematics" — large-scale, decentralized collaboration between humans and machines in which humans handle creative direction and AI systems handle technical execution at scale.
What this means
The Lean 4 formalizations make this a stronger claim than typical AI benchmark announcements: formal proofs can be checked independently of trust in OpenAI's reporting. That distinguishes this from marketing claims about benchmark scores that cannot be reproduced.
Still, key facts remain unverified or undisclosed. OpenAI has not published the prompts used, the number of failed attempts, or the total compute spent across all attempts (successful and unsuccessful). The $2,000-per-problem figure describes only the winning cases, which limits how much can be concluded about the model's overall cost-efficiency or success rate on hard problems.
The broader trend — AI labs publicly targeting long-standing open problems in mathematics and cryptography — suggests frontier labs are using unsolved research problems as a new benchmark category, one that's harder to game than standard test sets since the problems are, by definition, unsolved. Whether Astra's approach generalizes beyond these ten selected problems, or whether they were chosen because they were unusually tractable for an LLM-based approach, is not yet known.
Related Articles
Perplexity Says It Runs End-to-End Engineering Systems on OpenAI's GPT-6 Astra
Perplexity says it has shifted core engineering workflows, including code changes and production monitoring, onto OpenAI's GPT-6 Astra model. The claim comes from an OpenAI-published case study with no independent benchmark data released.
OpenAI Pauses New Pro Subscriptions as Astra Demand Overwhelms Infrastructure
OpenAI has temporarily disabled new sign-ups for its $200-per-month Pro plan, citing infrastructure strain from unprecedented demand for its Astra model. API, Go, and Plus plans remain unaffected.
OpenAI's GPT-6 Astra Beats Claude Fable 5.1 Nearly 3-to-1 in Autonomous Business Benchmark, Tops Drone Navigation Tests
Independent testing lab Andon Labs found OpenAI's GPT-6 Astra nearly triples Claude Fable 5.1's performance running a simulated vending machine business, averaging $15,515 versus $5,422. Astra also became the first model to beat human-AI baseline performance across all five Drone-Bench subtasks, including autonomous person-tracking via drone.
GPT-6 Astra Beats Ai2's MolmoAct2 on New Robotics Benchmark, Researcher Calls It a 'Step Change'
A new robotics benchmark called StationeryBench shows OpenAI's GPT-6 Astra completing 7 of 100 desk-object manipulation tasks versus zero for Ai2's MolmoAct2, with a median progress score of 46 against 12. Cornell/DeepMind researcher Yoav Artzi calls the result a 'step change in spatial reasoning.'
Comments
Loading...