OpenAI claims reasoning model disproved 80-year-old Erdős conjecture in geometry
OpenAI claims its new reasoning model has produced an original mathematical proof disproving a geometry conjecture first posed by Paul Erdős in 1946. The company says this is the first time AI has autonomously solved a prominent open problem central to a field of mathematics, with verification from mathematicians including Thomas Bloom and Noga Alon.
OpenAI Claims Reasoning Model Disproved 80-Year-Old Erdős Conjecture
OpenAI says its new general-purpose reasoning model has produced an original mathematical proof disproving a geometry conjecture first posed by mathematician Paul Erdős in 1946.
According to OpenAI, the model discovered "an entirely new family of constructions" that outperform what mathematicians believed were the best possible solutions for nearly 80 years. The company claims this marks "the first time AI has autonomously solved a prominent open problem central to a field of mathematics."
Verification and Context
Unlike OpenAI's previous claim in October 2025—when former VP Kevin Weil incorrectly stated GPT-5 had solved 10 Erdős problems, only to discover the solutions already existed in literature—the company this time published supporting statements from multiple mathematicians.
Verification came from:
- Noga Alon
- Melanie Wood
- Thomas Bloom, who maintains the Erdős Problems website
Bloom, who previously called Weil's October post "a dramatic misrepresentation," stated: "AI is helping us to more fully explore the cathedral of mathematics we have built over the centuries."
Technical Significance
OpenAI emphasizes the proof came from a general-purpose reasoning model, not a system specifically designed for mathematical problems. The company says this demonstrates AI systems can now "hold together long, difficult chains of reasoning and connect ideas across fields in ways researchers may not have previously explored."
The specific Erdős problem, model name, benchmark performance, and technical details of the proof were not disclosed in the announcement.
What This Means
If verified through peer review, this would represent a significant milestone in AI-assisted mathematical research—moving from pattern matching existing solutions to genuine novel discovery. The claim's credibility is strengthened by mathematician endorsements and OpenAI's apparent caution after last year's embarrassment. However, the lack of technical details, model specifications, and peer-reviewed publication leaves key questions unanswered. The broader implication: general reasoning models may now be capable of autonomous discovery in physics, biology, and engineering, not just mathematics.
Related Articles
OpenAI's GPT-5.6 Sol Adds Five Reasoning Effort Settings, Follows DeepSeep-R1 RLVR Training Method
OpenAI released GPT-5.6 Sol, a new reasoning model family that comes in three sizes with roughly five to six reasoning-effort settings each. The release follows the DeepSeek-R1 methodology of using reinforcement learning with verifiable rewards (RLVR), nearly two years after OpenAI's original o1 model popularized LLM-based reasoning.
OpenAI GPT-5.6 Sol, Terra, and Luna launch on Amazon Bedrock with 80-point Coding Agent Index score
OpenAI's GPT-5.6 model family is now generally available on Amazon Bedrock, introducing a three-tier system: Sol (flagship reasoning), Terra (balanced production), and Luna (fast inference). According to OpenAI, Sol scores 80 points on the Artificial Analysis Coding Agent Index and 73.5% on ExploitBench, establishing new benchmarks while using less than half the output tokens of competing models.
OpenAI restores chat sidebar in Mac app after user backlash over confusing redesign
OpenAI has updated its ChatGPT Mac app to restore direct access to chat conversations through a prominent sidebar toggle. The fix addresses user complaints following a July 10 redesign that replaced the native Mac client with an Electron-based app and buried the standard chat interface behind Work and Codex features.
Cline v4.0.9 Adds GPT-5.6 ChatGPT Models, Fixes Token Count Over-Reporting
Cline, the AI coding assistant VS Code extension, released v4.0.9 on July 16, 2024, adding support for GPT-5.6 ChatGPT subscription models. The update fixes a bug where token counts were over-reported from OpenAI-compatible providers due to improper handling of cumulative usage snapshots.
Comments
Loading...