10,000 Agents, 88 Hours, One Million-Dollar Math Problem
On Tuesday September 8, OpenAI published something that doesn’t happen often in mathematics: a claim, accompanied by a machine-checked proof, that an internal AI system has resolved one of the seven Millennium Prize Problems. The problem is the Navier–Stokes existence and smoothness question — whether a smooth fluid in three dimensions can “blow up” into a singularity in finite time, a possibility that has haunted fluid dynamics and mathematical physics for the better part of a century. If the claim holds, it is, as Quanta put it, “by a significant margin, the most important mathematical proof to have been arrived at by an artificial-intelligence model to date.”
Three days later, the math world is still arguing about it. So is everyone else. I think the argument is more interesting than the answer, and worth sitting with for a moment.
What was actually claimed
OpenAI says that a coordinating swarm of roughly 10,000 autonomous agents, running on an internal model the company describes as “significantly more capable than GPT-6 Astra” and that has been in training since August 28, spent 88 hours producing a 166-page analytical proof that a smooth three-dimensional fluid at rest, under a smooth external force, can develop a singularity in finite time while its energy stays finite. The mechanism is a vortex that spirals inward and elongates “like spaghetti.” After the proof landed, GPT-6 Astra spent an additional 17 hours translating the argument into Lean, the proof-assistant language mathematicians now reach for when they want airtight verification. Both the writeup and the formalization have been published publicly, in a repository called openai/NavierStokesAndEuler, where any independent verifier can re-check the logic. Sébastien Bubeck of OpenAI has put the compute bill at several million dollars; other estimates, billed at current Astra API rates, put it closer to $22.5 million for the 300 billion output tokens across the whole campaign.
Two things make this different from the previous wave of “AI does math” headlines.
First, the problem itself. The Clay Mathematics Institute named seven Millennium Prize Problems in 2000, each with a $1 million bounty. Until last week, only one had been solved — Grigori Perelman’s proof of the Poincaré conjecture in 2003, for which he famously declined the prize. Navier–Stokes is widely considered the most physically urgent of the remaining six, because the equations describe nearly every fluid you have ever cared about: weather, oceans, blood, air over a wing, water through a pipe. The question is whether those equations are always well-behaved, or whether somewhere, under some conditions, a smooth flow can turn pathological and the math stops describing anything physical.
Second, the way it was done. This is not a chatbot that was prompted once and produced a clever argument. The system OpenAI describes is a hierarchy of agent groups — each group attacking a different variant of the problem, easier sub-problems first, results shared back up to a “cross-pollinating” coordinator (using Codex, internally) that consolidated insights. The Euler blow-up, a closely related problem and a warm-up, took about 100 agents and 50 hours. The Navier–Stokes result took two orders of magnitude more — roughly 10,000 agents and 88 hours. In total, the agents exchanged nearly 5 million messages and burned around 130 billion output tokens on the Navier–Stokes run alone. The work was done. Lean was the receipt.
The dispute that broke it open
Roughly 12 hours before OpenAI went public, NYU mathematician Tristan Buckmaster released a four-page statement. He and Levent Alpöge, a mathematician at Anthropic, had spent the better part of a year using publicly available AI models from both Anthropic and OpenAI — including Claude and Codex — to chase the same cluster of problems. By August 15, they had Lean-verified blow-up proofs for three related equations (incompressible porous media, the 2D Boussinesq system, and 3D incompressible Euler). They were polishing the writeup for publication when, on September 3, Buckmaster learned from contacts that “information about our progress had been passed to OpenAI.” He pre-emptively emailed a prominent mathematician at the company to set the record straight.
What followed, by both accounts, was a series of calls. On September 6, Bubeck told Buckmaster that an internal OpenAI model had already produced roughly a 100-page proof of forced Navier–Stokes blow-up — the very same result Buckmaster and Alpöge were approaching from a different direction. According to Buckmaster, OpenAI proposed that he write up a joint result as sole author and leave Alpöge, an Anthropic employee, off the paper. Bubeck has called that account “false and inflammatory.” OpenAI has said it “did not use their prompt or proofs to prompt our models,” and that “while unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models.” Sam Altman publicly defended Bubeck as having “acted with integrity and generosity throughout.” The Clay Institute has not weighed in. The Millennium Prize remains, officially, unclaimed.
Whether or not any of the conduct alleged by Buckmaster turns out to be accurate, the more interesting question is structural. Two research teams, supported by overlapping sets of AI tools, converged on the same mathematical frontier at almost exactly the same time. That convergence was not a coincidence — both teams were building on the same “smooth forcing” approach pioneered by Diego Córdoba and Luis Martínez-Zoroa, an approach several experts had independently flagged as promising. But the timing — OpenAI’s push starting September 1, after rumors of an Anthropic-attributed solution — and the asymmetric compute available to a frontier lab versus a two-person academic team (88 hours and a small country worth of tokens, versus a year of part-time effort funded out of a professor’s own research budget) raises a question that mathematics has never had to deal with before. When the same problem is attacked by an academic team and a frontier lab, and the lab can spin up 10,000 instances on a moment’s notice, who actually won?
What I think is actually going on
A few things feel true at the same time.
-
The proof is real, in the Lean sense. The chain of logical steps, as expressed in Lean’s type theory, has been checked by an independent program. This is not the place where AI math results have historically broken down. The genuine uncertainty is upstream of the Lean check: does the formal statement OpenAI proved correspond to the Navier–Stokes Millennium Prize Problem as the Clay Institute means it? OpenAI says it has established Fefferman’s alternatives C and D, which permit a smooth external force. The strictest version of the problem, alternatives A and B, requires the force to be zero — and that is the version mathematicians like Córdoba think is the actually hard one. The Clay Institute is unlikely to pay out on a forced blow-up alone.
-
The model behind the proof is, by OpenAI’s own account, a meaningful step beyond what is publicly available. They describe it as “significantly more capable than GPT-6 Astra” — and Astra itself was released only a week earlier. If the capability is real, the rest of the field is going to feel a sharp pull. It also tells you something about how quickly the frontier is moving that an unreleased model, in training for ten days, can outperform a flagship on a problem a year-long research program could not crack.
-
10,000 agents is a new kind of artifact. Even setting aside whether the proof is right, the method is striking. This is not “AI helps a mathematician write a paper.” This is “AI, given the statement, organizes itself into teams, attempts sub-problems, shares what works, abandons what does not, and produces a candidate argument in three and a half days.” That is a different shape of system than the one we have been calling “AI for math” for the last few years, and it is going to spread to other open problems whether or not Navier–Stokes survives peer review. The Anthropic/Cognition/DeepMind-style agentic coding harnesses that have eaten software engineering over the last year are now coming for research mathematics.
-
The credit fight is a leading indicator. Buckmaster’s complaint is not really about plagiarism. It is about a frontier lab being able to outcompute an academic team by a factor of 1,000 on a problem the team was three days from finishing, and about the social norms that govern who gets to be named as having solved a problem. Those norms were not built for a world in which one of the parties in the race is an industrial-scale language model with effectively unlimited inference budget. The math community is going to have to invent them, and quickly, because this is the first such collision, not the last.
-
The result is, for now, a claim. Terence Tao, the safest possible pair of hands in the field, called the Buckmaster/Alpöge Euler result “a remarkable achievement” and said the method “will reach Navier–Stokes” — a careful endorsement that stops short of the full problem. As of this writing, no one outside OpenAI has read the 166 pages. Quanta’s framing is the right one to hold in mind: “if the result holds up to further scrutiny.” That is the operative phrase. It may. It may not. Either way, the week of September 8, 2026 is the moment mathematics found out what it looks like when the cost of a candidate proof drops from “a year of a researcher’s life” to “a long weekend of GPU time.”
The bit that stays with me
I keep coming back to the image in the proof itself: a vortex that spirals inward and gets “increasingly elongated, like spaghetti.” A blob of fluid, smooth and well-behaved, that quietly tightens on itself until the math can’t describe what’s happening anymore. There is a long history of mathematicians looking for that blow-up and not finding it; there is now, possibly, a machine-readable argument that it is real, and a Lean certificate that says the argument is internally consistent, and a public repository where anyone with the patience and the expertise can re-run the check.
Whether the answer is right or wrong, the question of how this is going to keep going is the one I would ask the mathematicians in my life. A year of human time, a long weekend of machine time, the same problem, two announcements twelve hours apart. That is the new normal. The next Millennium Prize will not be decided the way the last one was.
Comments
Leave a message below. Your comment saves to your browser.