analysisOpenAI

Mathematicians' group calls for OpenAI boycott after release of 700+ AI-generated proof files

TL;DR

The Association for Human Mathematics (AHM), chaired by Fields Medalist Terence Tao, is urging mathematicians to stop working with OpenAI after the company released more than 700 AI-generated manuscripts at once. The group says the release violates scientific norms. Critics say many of the papers are too dense to verify without AI assistance.

4 min read
0

The Association for Human Mathematics (AHM) is urging mathematicians to stop working with OpenAI after the company published more than 700 AI-generated mathematical manuscripts in a single release, according to reporting by The Decoder. The group accuses OpenAI of violating core norms of scientific research.

Fields Medalist Terence Tao, who chairs the AHM, shared the statement as a guest post on his blog. The statement says: "Mathematicians did not ask for this work to be done." It calls the release "not a demonstration of scholarship, but a demonstration of power."

What OpenAI released

  • Volume: More than 700 files, published at once.
  • Scope: Results span complexity theory, number theory and algorithm research, including claimed partial progress on several remaining Clay Millennium Problems. Complexity theorist Scott Aaronson says any one of them would have ranked among the year's top results on its own.
  • Readability: Many papers are reportedly so dense and oddly written that even leading experts need AI help to decipher them, if they can at all.
  • Unique Games Conjecture: The release includes a claimed proof. Dana Moshkovitz, who has worked on the problem for her entire career, called the paper "horribly written" and said the proof relies on "some alien craziness," according to Aaronson.

None of these results has been independently verified. The model used, its parameter count, context window, pricing and training cutoff have not been disclosed. The Decoder reports the model was likely OpenAI's current internal one, which could reach paying ChatGPT customers in the coming months.

Compute figures

Aaronson reports, citing unnamed sources, that about 8,000 problems were tested overall. The success rate was roughly 5%, using an average of three hours of compute at GPT-Pro level per problem. OpenAI has not confirmed these figures in the source material. Aaronson also reports that AI companies are quietly probing weaknesses in cryptographic protocols, an area absent from the release.

Background to the dispute

The AHM statement escalates a months-long debate. OpenAI previously claimed an internal model solved more than 100 open math problems in a single month, including an approach to the Navier-Stokes Millennium Problem. It then set up an advisory group, AGMAI, at the Institute for Advanced Study. The group has no say over the pace of OpenAI's internal research.

The AHM says OpenAI ignored AGMAI's central premise that advanced math problems shouldn't be tested on internal models. The statement also references OpenAI's copyright lawsuits, implying the results rely on training on mathematicians' work.

Separately, OpenAI announced a Navier-Stokes solution shortly before two mathematicians were due to present their own AI-assisted solution. The two had used ChatGPT. OpenAI denied using data from those interactions to train its system.

Tao previously joined 24 other Fields Medalists, including Peter Scholze, Maryna Viazovska and Martin Hairer, in warning of a "severe misalignment" between AI industry goals and those of mathematics. On Mastodon, he argues that solved problems cannot be "unsolved," and that knowing a solution exists can "contaminate" the search for other approaches. He proposes a "Math 2.0" era that values explanation and community-building over problem-solving.

Two release models

Aaronson contrasts OpenAI's batch release with what he calls the "Anthropic model." Anthropic partnered with two algorithm researchers whose AI model supplied the key idea for disproving two decades-old conjectures. The researchers were compensated and wrote a human-readable proof. Each approach has drawbacks. OpenAI's leaves the community to make proofs readable for free. Anthropic's lets a private company choose which mathematicians act as "emissaries."

Mathematicians are divided. Some commenters on Tao's blog back the boycott and call for better protection of preprint servers like arXiv against mass training-data collection. Others call the AHM's position unrealistic.

What this means

The dispute is no longer only about whether AI can solve open problems. It is about what counts as "solved" when verification is the bottleneck. Output at this scale shifts the cost to human reviewers, and no norm yet governs who bears it.

At a reported 5% success rate and roughly three hours of compute per attempt, mass problem-solving looks cheap relative to the human effort of checking the results. That asymmetry will likely push journals, preprint servers and funding bodies toward formal verification requirements or disclosure rules. A boycott is unlikely to slow OpenAI's math work on its own. But the Anthropic comparison suggests labs will face pressure to pair releases with human-written explanations and compensated collaborators.

Related Articles

analysis

Only 10 of OpenAI's 719 math manuscripts include chain of thought, falling short of expert guidelines

OpenAI released 719 manuscripts claiming solutions to open math problems, but only 10 include the model's chain of thought. A Cambridge and King's College London paper also documents at least two discrepancies between the natural-language and Lean versions of OpenAI's Navier-Stokes-derived result.

research

OpenAI publishes 372 AI-generated math results on GitHub, claims they solve or advance open problems

OpenAI has published 372 mathematical results generated by an unnamed internal frontier model, hosted on GitHub instead of in peer-reviewed journals. The company claims each result solves or substantially advances an open problem, at an average of about three hours of ChatGPT Pro Thinking compute per result. Many include Lean formalizations, but independent validation of significance is still pending.

product update

OpenAI to watermark ChatGPT and Codex text in the EU with textGrain; API watermarking is opt-in worldwide

OpenAI will switch on invisible text watermarks called textGrain for ChatGPT and Codex users in the EU over the coming weeks. API watermarking will be opt-in worldwide, unlike Anthropic's mandatory approach for Claude. OpenAI's own data shows detection drops sharply when text is edited.

analysis

Ramp AI Index: US business AI spending falls while usage rises about 50% from July peak

US companies are spending less on AI even as usage hit a record high at the end of September, according to the latest Ramp AI Index. Ramp economist Ara Kharazian attributes the drop almost entirely to price competition between OpenAI and Anthropic. In the last week of September, Anthropic took 51% of token spending and OpenAI 44.5%.

Comments

Loading...