On technical spectacle alone, OpenAI's announcement is difficult to compete with.
The company says an internal model still in training and significantly more capable than GPT-6 Astra produced a proposed solution to the Navier–Stokes existence and smoothness problem.
Navier–Stokes is one of the seven Millennium Prize Problems designated by the Clay Mathematics Institute, each historically associated with a $1 million award for a recognized solution.
OpenAI says it does not intend to claim the prize.
This was not one brilliant AI agent
The system was built around large groups of coordinating agents with access to tools, cached internet material and executable code.
Different groups received different formulations of the problem so they could explore competing approaches in parallel.
The group associated with the Navier–Stokes result involved on the order of 10,000 concurrent agents.
OpenAI also used Codex to consolidate promising ideas from different groups and feed those insights back into additional searches.
It is closer to a synthetic research institute operating at extreme parallelism than a single chatbot writing a proof.
The agents used 130 billion output tokens
OpenAI launched the first agents on September 1.
The company says the system reached its Navier–Stokes result on September 5, approximately 88 hours later.
Lean formalization and verification then took another 17 hours using GPT-6 Astra.
For Navier–Stokes alone, the agents exchanged about 2.7 million messages and generated approximately 130 billion output tokens.
Across all of the mathematical problems attempted during the broader experiment, OpenAI reports roughly 300 billion output tokens.
The million-dollar prize is almost the small number
Outside estimates have attempted to price 130 billion tokens using commercial-equivalent model costs and reached figures substantially above the $1 million Clay prize.
That calculation should not be taken literally.
OpenAI is not paying its own internal inference bill at public API prices, and the model involved is not a commercial product with a published token price.
Still, the comparison captures something about the economics of AI-driven research: enormous computational resources can now be concentrated on a scientific objective whose traditional reward looks small by comparison.
A rumor is what started the sprint
OpenAI openly acknowledges that timing mattered.
On September 1, its researchers heard a rumor that two Millennium Prize Problems had been resolved. That rumor encouraged them to test the new internal model across the remaining open problems.
OpenAI later concluded that the information was connected to work by NYU mathematician Tristan Buckmaster and Levent Alpöge, a mathematician employed by Anthropic.
That is where the research announcement turned into an academic dispute.
Buckmaster and Alpöge were already exploring related mathematics
The pair had been pursuing singularity-formation results in fluid equations and were themselves making extensive use of AI tools, including Codex and Anthropic systems.
Their result was not identical to OpenAI's.
Buckmaster and Alpöge produced a result concerning forced Euler equations. OpenAI says its agents independently obtained an unforced Euler result and then used that work as a stepping stone toward Navier–Stokes.
The disagreement is therefore not simply an allegation that OpenAI copied the same theorem.
It is about how a promising research direction became known, how quickly OpenAI concentrated resources on it and who should receive scientific credit for identifying the route.
Buckmaster questioned how OpenAI arrived there so quickly
Buckmaster publicly described his conversations with OpenAI and raised concerns about his previous use of Codex.
His work had involved entering drafts, mathematical arguments and developing ideas into AI systems. When OpenAI rapidly arrived at a result in an unusually similar area, he asked whether those interactions could have influenced the internal system.
That question became one of the central controversies surrounding the announcement.
OpenAI says its researchers and agents did not see Buckmaster and Alpöge's unpublished work before it became public.
OpenAI says it investigated the Codex prompts
On September 10, OpenAI added a more specific statement to its publication.
The company says an investigation established that Buckmaster's Codex prompts from the previous two months could not have influenced the system that generated the result, including through training.
OpenAI says the internal model was created through large-scale reinforcement learning on top of an earlier pretrained model.
That addresses the specific allegation more directly than simply saying nobody manually opened a researcher's conversations.
It leaves a broader question intact: what separation should scientists expect between the proprietary AI tools they use for research and the internal research systems operated by the same company?
The publication negotiations made the situation messier
According to Buckmaster's account reported by multiple outlets, OpenAI subsequently discussed different options for coordinating publication and credit.
One possibility would have given Buckmaster a prominent role in presenting OpenAI's result while leaving Alpöge out, with his employment at Anthropic creating an obvious corporate complication.
Buckmaster rejected the arrangement.
OpenAI disputes the characterization that it was attempting to take credit improperly and says it sought to recognize Buckmaster and Alpöge's priority on their separate forced-Euler result.
It is a remarkably 2026 scientific dispute: two mathematicians collaborate, one works at Anthropic, both use frontier AI tools, and a fluid-dynamics problem ends up entangled with competition between AI labs.
No, Clay has not handed over the million dollars
The most important caveat is also the easiest one to lose in the headlines.
OpenAI has published a proposed solution with a Lean formalization. The Millennium Prize Problem is not yet officially resolved under Clay Mathematics Institute rules.
Before Clay will even consider a proposed solution for the prize, it must appear in a qualifying publication, at least two years must pass, and the result must achieve general acceptance within the global mathematics community.
That two-year scrutiny period is a minimum requirement.
It cannot be parallelized into a weekend by adding another few thousand agents.
Lean verification is powerful, but it is not scientific consensus
The formal proof is nevertheless a major part of why the announcement is being taken seriously.
A proof assistant such as Lean can mechanically check that a formalized argument follows the logical rules encoded in its environment.
That removes a large class of subtle human proof errors.
What Lean does not do is independently determine that every formal assumption perfectly corresponds to the official problem statement, that the formalization faithfully captures every mathematical claim surrounding it, or that the result satisfies Clay's institutional criteria.
Machine verification is an exceptionally strong piece of evidence. It is not a digital button labeled “collect $1 million.”
The bigger shock may be what this does to unpublished ideas
The experiment demonstrates what happens when a laboratory can concentrate thousands of capable agents on one research direction within hours.
A human researcher may spend years identifying which representation of a problem is promising.
A sufficiently large agent system can then test enormous numbers of branches from that insight in parallel.
That changes the value of rumors, partial drafts and unpublished intuition.
The complete proof may no longer be the only scarce scientific object. Simply knowing which route is worth attacking can become extremely valuable.
Science itself is starting to look like the benchmark
OpenAI frames the result primarily as evidence of rapid model progress rather than an attempt to win a mathematics prize.
Viewed that way, the demonstration is difficult to dismiss.
Ten thousand agents, millions of messages, nine-figure token usage and a few days of concentrated work were enough to produce a serious proposed answer to a problem that has resisted mathematicians for decades.
Even if the proof ultimately requires corrections, the research method has already changed the discussion.
Academic norms around priority, collaboration and credit were built for researchers operating at human speed.
OpenAI has just shown what happens when someone adds 10,000 synthetic colleagues to the project.