OpenAI publishes 372 AI-generated math results on GitHub, claims they solve or advance open problems
OpenAI has published 372 mathematical results generated by an unnamed internal frontier model, hosted on GitHub instead of in peer-reviewed journals. The company claims each result solves or substantially advances an open problem, at an average of about three hours of ChatGPT Pro Thinking compute per result. Many include Lean formalizations, but independent validation of significance is still pending.
OpenAI has published 372 mathematical results generated by an internal frontier model, hosting them in a GitHub repository rather than submitting them to academic journals. The company claims each result solves an open problem or makes substantial progress toward one. The model has not been named or released.
What OpenAI published
According to OpenAI, the collection includes improvements to major computer algorithms and advances related to the Riemann hypothesis. The repository includes revision logs and citations. OpenAI has not disclosed the model's name, parameter count, context window, or pricing.
OpenAI says nearly every result came from a single prompt to a single AI agent, though some required multiple attempts. The company says each result consumed roughly three hours of ChatGPT Pro Thinking compute on average. OpenAI shared only average compute costs, not per-problem figures, and did not publish the prompts.
The company also published methodology details: summaries of the model's reasoning process, statistics on how many problems the model attempted, and compute cost estimates.
The Navier-Stokes comparison
OpenAI says the same model previously produced a solution to a Navier-Stokes problem that has been under formal review for weeks. According to the report, that solution required a swarm of 10,000 agents and millions of dollars in compute, far more than the single-agent runs behind most of the 372 results. The Navier-Stokes result has not been confirmed by outside reviewers.
Formal verification as a review strategy
Many of the proofs ship with formalizations in Lean, a language for machine-checkable proofs, and OpenAI plans to add more. The rationale is review capacity: the volume of AI-generated output could overwhelm manual checking by mathematicians.
Lean can confirm that a proof is logically valid. It cannot judge whether a result is relevant, original, or interesting. Those judgments remain with human mathematicians.
Advisory group and limits
OpenAI consulted the Advisory Group on Mathematics and Artificial Intelligence at the Institute for Advanced Study, which includes Fields Medal winner Timothy Gowers, and says it loosely followed the group's public recommendations. It set one boundary in advance: the mathematicians can advise on how results are communicated, but not on whether or how fast they are produced.
OpenAI plans to fund workshops and conferences on understanding AI-produced results. It acknowledges its citations and presentation need improvement. It also says it is working on a responsible release of the model to "directly empower scientists with state-of-the-art capabilities." No release date has been given.
Community reaction
Reactions among mathematicians range from interest to frustration. In an open letter titled "A Severe Misalignment of AI in Mathematics," 25 Fields Medal winners argued that problem-solving is a proxy for the real goal of conceptual understanding. They warned that mass-producing true statements could damage the field rather than generate new ideas.
Gowers has warned that within one to two decades, the mathematical literature could grow enormously while no human community truly understands it. Terence Tao has argued that training young mathematicians should emphasize the human side and tightly limit AI tool use.
What this means
The significant claim here is not any single theorem but the cost structure. If a typical result really takes about three hours of ChatGPT Pro Thinking compute, the bottleneck in mathematics shifts from generating results to evaluating them. That is the logic behind OpenAI's emphasis on Lean.
Several things limit how much weight the numbers deserve today. The model is unnamed and unavailable. Prompts are unpublished and per-problem costs are undisclosed, so the "three hours" figure cannot be independently checked. The count of 372 does not show how many results are substantive. The attempt statistics OpenAI published should help with that, and the repository's contents can now be scrutinized.
Bypassing journals is a deliberate choice, and it sets up a conflict over who certifies mathematical knowledge. Formal verification can establish correctness, but significance still depends on the community, which is already publicly divided. How mathematicians receive these results, and whether the Navier-Stokes review concludes, will test OpenAI's claims more than the repository's size does.
Related Articles
OpenAI ships opt-in textGrain text watermarking in API, with EU ChatGPT and Codex rollout to follow
OpenAI has launched textGrain, an invisible statistical text watermark, as opt-in for API customers worldwide on select models. ChatGPT and Codex output in the EU will be watermarked in the coming weeks, in response to the EU AI Act. OpenAI says detection drops from about 92% to 17% when 25% of words in a 400-token passage are replaced.
OpenAI to watermark ChatGPT and Codex text in the EU under AI Act; API opt-in available worldwide
OpenAI will add an invisible watermark to text generated by ChatGPT and Codex in the European Union to comply with the EU AI Act's transparency rules. Developers anywhere can enable it on select API models starting today, but it is off by default. OpenAI's own tests show detection falling from about 92% to 66% after 10% of words are replaced with synonyms.
OpenAI to watermark ChatGPT and Codex text in the EU with textGrain; API watermarking is opt-in worldwide
OpenAI will switch on invisible text watermarks called textGrain for ChatGPT and Codex users in the EU over the coming weeks. API watermarking will be opt-in worldwide, unlike Anthropic's mandatory approach for Claude. OpenAI's own data shows detection drops sharply when text is edited.
OpenAI publishes startup guide for GPT-6 family covering model choice, reasoning effort and tool coordination
OpenAI has published "A model guide for the GPT-6 family," a practical guide aimed at startups. It covers choosing GPT-6 models, tuning reasoning effort, improving prompts and skills, coordinating tools, and preparing workflows for production. The summary gives no pricing, context window or benchmark figures.
Comments
Loading...