OpenAI announced on 8 September 2026 that a collection of roughly 10,000 AI agents, running an unreleased model, arrived at a solution to the Navier–Stokes Millennium Problem within 88 hours. The Navier–Stokes equations describe fluid flow and the unsettled theoretical question asks whether smooth initial conditions can develop singularities (for example, velocities blowing up to infinity) in finite time.
OpenAI says the effort began 1 September after hearing others were close. The company acknowledged competitors — notably NYU's Tristan Buckmaster and Levent Alpöge of Anthropic — had worked on the problem, sometimes using OpenAI's Codex tool. Buckmaster publicly questioned whether his drafts or private Codex sessions influenced OpenAI's model. OpenAI initially said it could not rule that out, then later stated an investigation found those prompts could not have affected the result.
OpenAI's published proof was checked in Lean, a formal proof assistant that verifies each logical step. That makes the underlying argument highly likely to be correct in a formal sense. Formal verification provides a referee that is immune to rhetorical influence: if Lean accepts a proof, the formal steps follow the specified axioms and rules.
Formal correctness is not the same as human comprehension. Terence Tao compared automated extraction of solutions to using heavy machinery at an archaeological dig: you can retrieve treasure but destroy the contextual detail that gives it meaning. Tao suggested some problems might need informal protections against automated solutions—akin to a spoiler norm—so that human discovery and learning remain possible.
Historically only one Millennium Problem had been solved before: Grigori Perelman's proof of the Poincaré Conjecture. Perelman refused the prize money in 2010, citing others' contributions. OpenAI has said it will not claim the prize either. Media and commentators noted the contrast between Perelman's stance and OpenAI's rapid, compute-intensive run: one mathematician declined the reward on grounds of credit, while a tech company spent significant resources after picking up a rumor on social media.
Both teams referenced a strategy developed by Spanish mathematicians Diego Córdoba and Luis Martínez-Zoroa. Princeton mathematician Charles Fefferman credited Córdoba and Martínez-Zoroa's approach as central to the recent progress. The broader picture is that human ideas guided the search strategy even as automation scaled the execution.
The episode highlights tensions in mathematics and adjacent fields as automated systems scale problem-solving. Formal verification tools like Lean can confirm correctness, but the community now faces choices about norms for disclosure, credit, and whether some research avenues should be insulated to preserve human-driven discovery and pedagogy. The episode also raises trust questions about provenance when powerful models are trained or run on materials that include private human work.