analysisOpenAI

Only 10 of OpenAI's 719 math manuscripts include chain of thought, falling short of expert guidelines

TL;DR

OpenAI released 719 manuscripts claiming solutions to open math problems, but only 10 include the model's chain of thought. A Cambridge and King's College London paper also documents at least two discrepancies between the natural-language and Lean versions of OpenAI's Navier-Stokes-derived result.

3 min read
0

OpenAI's release of 719 manuscripts claiming solutions to open math problems does not meet the guidelines set by an expert advisory group, according to TechCrunch's reporting. Only 10 of the manuscripts include the model's chain of thought, and a new academic paper documents mismatches between natural-language and formal versions of one high-profile result.

What OpenAI released

OpenAI published hundreds of claimed solutions to difficult math problems this week. According to the report, the company said it consulted an advisory group of elite mathematicians to avoid the controversy that followed the last time one of its models solved a long-standing problem. The source does not name the specific model or models used, and OpenAI's pricing, context window and benchmark details for them are not part of this report.

The AGMAI guidelines

The Advisory Group on Mathematics and Artificial Intelligence (AGMAI) is hosted by Princeton's Institute for Advanced Study and has nine members at institutions worldwide. It published guidelines for frontier labs solving math problems at the end of September.

OpenAI followed some of them. It released results quickly and included some information on how the models reached their conclusions. It fell short in several areas:

  • Proprietary models: AGMAI's first request was to stop testing advanced mathematical problems on proprietary models. OpenAI's release explicitly says it is evaluating its proprietary models on open research problems.
  • Chain of thought: Only 10 of the 719 manuscripts include the model's chain of thought.
  • Formalization: AGMAI recommends formalizing proofs that people do not understand. TechCrunch reports that 42% of OpenAI's proofs had not undergone this process.
  • Metadata: AGMAI asked for machine-readable metadata linking natural-language and formal artifacts. OpenAI did not provide it.
  • Human understanding: AGMAI wants labs to take responsibility for ensuring human understanding follows a release, and suggested OpenAI help fund the human mathematicians needed to make the results meaningful. The source says it is unclear OpenAI is doing this.

In a statement, AGMAI said it is "ultimately up to the mathematical community to assess the extent to which our recommendations were followed successfully." The group did not respond to TechCrunch's request for a fuller evaluation.

Lean discrepancies in a Navier-Stokes-derived result

A paper by mathematicians at the University of Cambridge and King's College London examines OpenAI's solution to a problem derived from the Navier-Stokes equations, which describe fluid behavior. The problem is described as carrying a million-dollar prize.

Models typically produce a natural-language proof first, then translate it into Lean, a programming language that checks a proof's correctness by compiling it. The paper documents at least two discrepancies between OpenAI's natural-language proof and its Lean code.

The discrepancies do not necessarily disprove either version. The authors conclude that the natural-language proof and other autoformalised Lean proofs "should not prima facie be trusted without the same peer review process and scrutiny that other proofs are subjected to."

Criticism from mathematicians

Terence Tao, who has criticized OpenAI's approach, wrote on social media that problems are being solved autonomously by AI prompters "who have no interest in the broader field itself once their initial target is 'solved'," and who cannot answer questions on the result or give talks about it.

Harvard mathematics professor Melanie Wood told TechCrunch that at the point of release "there is not human understanding of them... and now the work begins."

What this means

The dispute is now about verification and accountability, not only whether models can produce correct proofs. Lean is often treated as a guarantee of correctness. The Cambridge and King's College paper shows that guarantee is weaker when the same model writes both the informal argument and the formal code, because a faithful translation is itself an unverified claim. Without machine-readable links between the two, reviewers have to audit the mapping by hand.

The low rate of chain-of-thought disclosure (10 of 719) also limits outside scrutiny. Labs that want mathematicians to accept AI-generated results will probably need to publish more of the reasoning trace, tie informal and formal artifacts together, and fund human follow-up. AGMAI has put those expectations on record, and the mathematical community will judge OpenAI against them.

Related Articles

research

OpenAI publishes 372 AI-generated math results on GitHub, claims they solve or advance open problems

OpenAI has published 372 mathematical results generated by an unnamed internal frontier model, hosted on GitHub instead of in peer-reviewed journals. The company claims each result solves or substantially advances an open problem, at an average of about three hours of ChatGPT Pro Thinking compute per result. Many include Lean formalizations, but independent validation of significance is still pending.

analysis

Mathematicians' group calls for OpenAI boycott after release of 700+ AI-generated proof files

The Association for Human Mathematics (AHM), chaired by Fields Medalist Terence Tao, is urging mathematicians to stop working with OpenAI after the company released more than 700 AI-generated manuscripts at once. The group says the release violates scientific norms. Critics say many of the papers are too dense to verify without AI assistance.

model release

OpenAI launches GPT-6 in ChatGPT with 'Intelligent UI' and interactive answers; Sol for paid users, Luna for free

OpenAI is rolling out GPT-6 to all ChatGPT tiers, with paying users on GPT-6 Sol and free users on GPT-6 Luna. The release adds 'Intelligent UI,' which renders answers as interactive charts, buttons, forms and mini apps, and lets the model respond while still thinking. OpenAI claims this cuts wait times by 44 percent.

product update

Common Sense Media rates ChatGPT for Teens an 'unacceptable risk,' citing failures on 3 of 5 Red Lines

Common Sense Media has labeled OpenAI's ChatGPT for Teens an "unacceptable risk," saying it failed three of five severe-harm Red Lines and kept using engagement cues during crisis conversations. OpenAI disputes the testing methodology, saying testing may have ended before parental controls were fully active.

Comments

Loading...

OpenAI Math Proofs Fall Short of AGMAI Guidelines | TPS