OpenAI Claims Resolution to Navier–Stokes Millennium Prize Problem Amid Priority Dispute
OpenAI claims its unreleased internal model resolved the Navier–Stokes existence and smoothness problem, one of seven $1 million Millennium Prize Problems. NYU professor Tristan Buckmaster disputes the timeline, alleging OpenAI moved after hearing rumors of his own near-year-long collaboration with Anthropic's Levent Alpöge.
OpenAI says an unreleased internal model produced a resolution to the Navier–Stokes existence and smoothness problem, one of the seven Millennium Prize Problems that has carried a $1,000,000 bounty from the Clay Mathematics Institute since May 24, 2000. The claim, published by OpenAI on September 8, 2026, is now entangled in a public dispute over research priority.
According to OpenAI's own account, the effort began after the company "heard rumors" on September 1 that two Millennium Prize problems had been solved. The company says it then tasked its internal model — inspired by what it calls a "step change in performance" — against all open Millennium Prize problems and several other high-impact problems. OpenAI claims its agents reached a resolution to Navier–Stokes on September 5, roughly 88 hours after the first agents launched, with Lean formalization and verification taking an additional 17 hours using a model OpenAI refers to as GPT-6 Astra.
OpenAI states that across all attempted problems, its agents sent 4.9 million messages and generated approximately 300 billion output tokens. The Navier–Stokes work alone reportedly consumed 2.7 million messages and roughly 130 billion output tokens. At public API pricing for GPT-6 Astra, 300 billion output tokens would cost approximately $15,000,000, though OpenAI has not disclosed the actual cost structure of the internal model used.
The rumors OpenAI cites reportedly originated from Tristan Buckmaster, an NYU mathematics professor, and Levent Alpöge, a mathematician employed by Anthropic. According to a statement published by Buckmaster, the pair had worked on the problem for nearly a year using Claude and Codex — primarily a model referred to as GPT-5.6 Sol — before reaching a breakthrough on August 15, 2026. Buckmaster says that after word of their result spread through the mathematical community, OpenAI contacted them, initially declining to specify when its own effort began. Buckmaster states that when he eventually pressed the point, OpenAI confirmed its work started only "in the past few days," after information about the NYU-Anthropic collaboration had already reached the company.
Buckmaster also says he asked OpenAI directly whether its model had been trained on or had access to the Codex sessions where he and Alpöge stored their drafts throughout the project. He states OpenAI told him the model did not look up user data but did not answer his follow-up question about training data.
OpenAI's public statement addresses this directly: "We (the researchers and the agents) did not see any of their work through any means until they released it publicly — in particular, no specific user data was accessed in order to solve this problem. While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models." OpenAI adds that its proof differs from Buckmaster and Alpöge's, including differing results in the Euler case (forced versus unforced).
OpenAI reportedly offered Buckmaster a choice: wait for OpenAI's release, or co-author a paper on OpenAI's result — but told him Alpöge would not be invited as a co-author because of Anthropic's competitive relationship with OpenAI.
None of the models involved — GPT-5.6 Sol, GPT-6 Astra, or the unnamed internal model OpenAI used for the Millennium Prize sweep — have been publicly released or given disclosed pricing or benchmark scores.
What this means
The technical achievement, if verified through Lean formalization, would mark a significant moment for AI-assisted mathematics regardless of the priority dispute. But the controversy raises a sharper question: once a rumor circulates that a hard problem has been solved, does that rumor alone trigger a computational arms race, with well-resourced labs deploying agents to reach the answer first? Commentators have already drawn a parallel to cybersecurity, where the mere rumor of an unpatched vulnerability can be enough to trigger automated exploit-finding. The unresolved question about whether de-identified user data from Codex or Claude sessions could ever surface in a competitor's model training remains the most consequential unknown here — not for this one dispute, but for anyone using AI tools to develop original, high-value intellectual work before it's published.
Related Articles
OpenAI Claims Unreleased Model Solved Navier-Stokes Problem, Faces Credit Dispute
OpenAI claims an internal, unreleased model solved part of the Navier-Stokes Millennium Prize problem using roughly 10,000 concurrent agents over 88 hours. NYU mathematician Tristan Buckmaster has questioned whether the effort drew on his unpublished research with an Anthropic researcher.
Unverified 'GPT-6 Astra' Reportedly Completes Portal Solo in Under 24 Hours, No Official OpenAI Confirmation
A developer named cozyblaze posted on X that a model called 'GPT-6 Astra' completed Portal from start to finish without human intervention in 23 hours 43 minutes. OpenAI has not confirmed the existence of a model by that name, and all details come from a single third-party account.
OpenAI Launches GPT-6 Astra With Half the Message Allowance of GPT-5.6 Sol
OpenAI has begun rolling out GPT-6 Astra to top-tier ChatGPT plans, the API, Azure, and AWS Bedrock. The model delivers roughly half the usage allowance of GPT-5.6 Sol across comparable plans, with Plus and Business users gaining access in the coming days.
Simon Willison's Pelican Benchmark Shows GPT-6 Astra Outperforming GPT-5.6 Sol at Every Reasoning Level
Developer Simon Willison ran his signature 'pelican riding a bicycle' SVG test on newly-accessed GPT-6 Astra across five reasoning levels, comparing results against GPT-5.6 Sol, Terra, and Luna. Even Astra's lowest reasoning setting reportedly beat every Sol output, though Astra costs roughly twice as much per token.
Comments
Loading...