OpenAI Model Disproves 78-Year-Old Erdos Conjecture, Triggering Mixed Reaction From Mathematicians
OpenAI published a counterexample disproving the Unit Distance Conjecture, a geometric graph theory problem open since 1946, in what many mathematicians call the most significant AI math result yet. Reactions range from Terence Tao's cautious optimism to Timothy Gowers describing 'mixed feelings' about having the rug pulled out from under him.
OpenAI published a counterexample disproving the Unit Distance Conjecture in May 2026, cracking a geometric graph theory problem that had remained open since 1946. The result is one of hundreds of open problems tied to Hungarian mathematician Paul Erdos, and according to multiple mathematicians, it is the most significant AI-driven math result to date.
The pace of follow-up work has been fast. One week after OpenAI's publication, human researchers adapted the underlying proof technique to disprove a separate major conjecture. Since then, according to the report, AI-assisted math results have arrived almost daily — new counterexamples, pattern detection, and machine-checkable proof conversions.
Epoch AI, which built the FrontierMath benchmark, recently announced a second solved problem in its "FrontierMath: Open Problems" track, which draws questions from major unsolved areas of mathematics. OpenAI's newer Astra model was introduced alongside ten solutions of varying difficulty, though it did not solve any of the six remaining Millennium Prize Problems, each carrying a $1 million prize from the Clay Mathematics Institute. OpenAI researcher Noam Brown has claimed that additional compute could eventually change that outcome.
Epoch AI's benchmark data shows AI has not yet solved a single problem in its two hardest tiers, "Major Advance" and "Breakthrough," indicating the technology's current progress is concentrated in more tractable problem categories such as graph theory rather than the field's deepest open questions.
Researchers split on what it means
Abhishek Saha, a mathematics professor at Queen Mary University of London, wrote on X that frontier models are "at least as good as a solid and indefatigable PhD student" in his research area after using GPT-5.5 Pro for a full day of routine work that previously took weeks. He described becoming "increasingly a conductor, rather than doubling up as the whole orchestra," and predicted the field will split between mathematicians who adapt and those who resist, comparing the latter to the folk hero John Henry.
Fields Medal winner Timothy Gowers offered a more conflicted account. He said GPT-5.6 Pro solved a problem on its first attempt on two occasions after he had spent considerable time on it himself, calling the experience "very strange and not particularly pleasant," while still welcoming the solutions. His larger concern, raised in connection with the Leiden Declaration on Artificial Intelligence and Mathematics — an effort backed by the International Mathematical Union and signed by more than 3,000 mathematicians — is the potential "destruction of mathematical culture" if fewer researchers develop deep expertise while the published literature expands beyond what any human community can fully understand.
Terence Tao, speaking at the 2026 International Congress of Mathematicians, compared the current period to the early 20th century foundational crisis triggered by paradoxes and incompleteness theorems. Tao warned that mathematics could shift from a scarcity of proofs to an overload of them, arriving faster than the community can review, and argued mathematicians will need new workflows to decide which results matter and how they should be validated and taught.
A Carnegie Mellon University team separately solved an open Ramsey theory problem in April 2026 by combining SAT solvers, language-model-generated code, and formal proof verification, a result they explicitly tied to a "golden age" of mathematics predicted by Gowers in 2000 — one he also warned would not last indefinitely.
What this means
The Unit Distance Conjecture result is a genuine milestone, but the surrounding benchmark data tells a more limited story: AI is clearing decades-old problems in specific subfields like graph theory while making no progress on the hardest tiers of open mathematics, including the Millennium Prize Problems. The more consequential shift may not be raw problem-solving but the emerging debate over process — how proofs get verified, credited, and absorbed into the field once machines can generate them faster than humans can check them. The Leiden Declaration's 3,000-plus signatories suggest the math community is moving toward governance rather than resistance, but figures like Gowers make clear that even researchers who welcome the results have not resolved what it costs the profession to get there.
Related Articles
OpenAI's GPT-6 Astra Beats Claude Fable 5.1 Nearly 3-to-1 in Autonomous Business Benchmark, Tops Drone Navigation Tests
Independent testing lab Andon Labs found OpenAI's GPT-6 Astra nearly triples Claude Fable 5.1's performance running a simulated vending machine business, averaging $15,515 versus $5,422. Astra also became the first model to beat human-AI baseline performance across all five Drone-Bench subtasks, including autonomous person-tracking via drone.
GPT-6 Astra Beats Ai2's MolmoAct2 on New Robotics Benchmark, Researcher Calls It a 'Step Change'
A new robotics benchmark called StationeryBench shows OpenAI's GPT-6 Astra completing 7 of 100 desk-object manipulation tasks versus zero for Ai2's MolmoAct2, with a median progress score of 46 against 12. Cornell/DeepMind researcher Yoav Artzi calls the result a 'step change in spatial reasoning.'
Perplexity Says It Runs End-to-End Engineering Systems on OpenAI's GPT-6 Astra
Perplexity says it has shifted core engineering workflows, including code changes and production monitoring, onto OpenAI's GPT-6 Astra model. The claim comes from an OpenAI-published case study with no independent benchmark data released.
AWS Benchmark: OpenAI's GPT-5.6 Luna Beats GPT-5.4 Mini on Cost-Per-Correct-Answer Despite Similar List Price
An AWS blog post using an open-source benchmarking harness finds that GPT-5.6 Luna, Terra, and Sol on Amazon Bedrock deliver lower cost-per-correct-answer than OpenAI's cost-optimized GPT-5.4 Mini and Nano, once accuracy, token efficiency, and agent turn counts are factored in. The analysis also cites a July 30, 2026 price cut of up to 80% for GPT-5.6 Luna on Amazon Bedrock.
Comments
Loading...