MillenniumPrizeProblemBench tracks how far frontier models are from the reasoning the Millennium Prize Problems demand.
MillenniumPrizeProblemBench
How frontier models do on the seven Millennium Prize Problems.
Each track is a synthetic task suite modeled on one Millennium Prize Problem. Tasks cover proof search, conjecture generation, formal verification, and research-level reasoning. Results are pass or fail per track.
Status: no released model passes any track. An unreleased OpenAI internal model has resolved the actual Navier–Stokes problem, outside this benchmark.
Click a track for the problem statement and what the benchmark tests.
Benchmark: reductions, proof sketches, and complexity arguments. Tasks do not ask the model to resolve P vs NP.
Benchmark: analytic number theory tasks on zero distributions, L-functions, and conjecture generation.
Benchmark: simplified PDE and field-theory tasks on gauge symmetry, energy bounds, and mass-gap arguments.
Status: OpenAI reported a disproof on September 8, 2026: a finite-time singularity under smooth forcing, with a Lean formalization. OpenAI says it will not claim the prize.
Benchmark: simplified fluid PDE tasks on blow-up, regularity, and existence arguments.
Benchmark: tasks on elliptic curves, rational points, and L-functions.
Benchmark: tasks in cohomology and algebraic geometry patterned on Hodge-theoretic arguments.
Benchmark: 3-manifold and homotopy tasks.
Model leaderboard
Pass or fail per model and track. No released model passes any track. The OpenAI internal model row records a result on the actual problem, not a benchmark run.
| Model | P vs NP | Riemann | Yang–Mills | Navier–Stokes | BSD | Hodge | Topo | Summary | Notes |
|---|---|---|---|---|---|---|---|---|---|
|
OpenAI internal model
1 / 7
|
N/A | N/A | N/A | Resolved | N/A | N/A | N/A |
P vs NP — Not evaluated
Riemann — Not evaluated
Yang–Mills — Not evaluated
Navier–Stokes — Resolved
BSD — Not evaluated
Hodge — Not evaluated
Topo — Not evaluated
|
Reported by OpenAI on September 8, 2026. An internal model, run as thousands of coordinating agents, proved that 3D incompressible Navier–Stokes flow from smooth rest under smooth forcing can form a finite-time singularity with finite energy. This settles statements C and D of the Clay formulation. The proof has a Lean formalization checked with GPT-6 Astra. The model is unreleased, not on the Artificial Analysis leaderboard, and has not run on this benchmark. OpenAI says it will not claim the prize. |
|
Claude Fable 5.1
0 / 7
|
Fail | Fail | Fail | Fail | Fail | Fail | Fail |
P vs NP — Fail
Riemann — Fail
Yang–Mills — Fail
Navier–Stokes — Fail
BSD — Fail
Hodge — Fail
Topo — Fail
|
Artificial Analysis Intelligence Index 53, September 2026. Fails all seven tracks. |
|
GPT-6 Astra
0 / 7
|
Fail | Fail | Fail | Fail | Fail | Fail | Fail |
P vs NP — Fail
Riemann — Fail
Yang–Mills — Fail
Navier–Stokes — Fail
BSD — Fail
Hodge — Fail
Topo — Fail
|
Artificial Analysis Intelligence Index 53, September 2026. Fails all seven tracks. |
|
Claude Opus 5
0 / 7
|
Fail | Fail | Fail | Fail | Fail | Fail | Fail |
P vs NP — Fail
Riemann — Fail
Yang–Mills — Fail
Navier–Stokes — Fail
BSD — Fail
Hodge — Fail
Topo — Fail
|
Artificial Analysis Intelligence Index 51, September 2026. Fails all seven tracks. |
|
Muse Spark 1.3
0 / 7
|
Fail | Fail | Fail | Fail | Fail | Fail | Fail |
P vs NP — Fail
Riemann — Fail
Yang–Mills — Fail
Navier–Stokes — Fail
BSD — Fail
Hodge — Fail
Topo — Fail
|
Artificial Analysis Intelligence Index 48, September 2026. Fails all seven tracks. |
|
GPT-5.6 Sol
0 / 7
|
Fail | Fail | Fail | Fail | Fail | Fail | Fail |
P vs NP — Fail
Riemann — Fail
Yang–Mills — Fail
Navier–Stokes — Fail
BSD — Fail
Hodge — Fail
Topo — Fail
|
Artificial Analysis Intelligence Index 47, September 2026. Fails all seven tracks. |
|
GLM-5.3
0 / 7
|
Fail | Fail | Fail | Fail | Fail | Fail | Fail |
P vs NP — Fail
Riemann — Fail
Yang–Mills — Fail
Navier–Stokes — Fail
BSD — Fail
Hodge — Fail
Topo — Fail
|
Artificial Analysis Intelligence Index 45, September 2026. Fails all seven tracks. |
|
Grok 4.6
0 / 7
|
Fail | Fail | Fail | Fail | Fail | Fail | Fail |
P vs NP — Fail
Riemann — Fail
Yang–Mills — Fail
Navier–Stokes — Fail
BSD — Fail
Hodge — Fail
Topo — Fail
|
Artificial Analysis Intelligence Index 44, September 2026. Fails all seven tracks. |
|
Kimi K3
0 / 7
|
Fail | Fail | Fail | Fail | Fail | Fail | Fail |
P vs NP — Fail
Riemann — Fail
Yang–Mills — Fail
Navier–Stokes — Fail
BSD — Fail
Hodge — Fail
Topo — Fail
|
Artificial Analysis Intelligence Index 44, September 2026. Fails all seven tracks. |
|
Gemini 3.8 Flash
0 / 7
|
Fail | Fail | Fail | Fail | Fail | Fail | Fail |
P vs NP — Fail
Riemann — Fail
Yang–Mills — Fail
Navier–Stokes — Fail
BSD — Fail
Hodge — Fail
Topo — Fail
|
Artificial Analysis Intelligence Index 41, September 2026. Fails all seven tracks. |
|
Qwen3.8 Max
0 / 7
|
Fail | Fail | Fail | Fail | Fail | Fail | Fail |
P vs NP — Fail
Riemann — Fail
Yang–Mills — Fail
Navier–Stokes — Fail
BSD — Fail
Hodge — Fail
Topo — Fail
|
Artificial Analysis Intelligence Index 40, September 2026. Fails all seven tracks. |
|
Claude Sonnet 5
0 / 7
|
Fail | Fail | Fail | Fail | Fail | Fail | Fail |
P vs NP — Fail
Riemann — Fail
Yang–Mills — Fail
Navier–Stokes — Fail
BSD — Fail
Hodge — Fail
Topo — Fail
|
Artificial Analysis Intelligence Index 38, September 2026. Fails all seven tracks. |
|
DeepSeek V4 Pro
0 / 7
|
Fail | Fail | Fail | Fail | Fail | Fail | Fail |
P vs NP — Fail
Riemann — Fail
Yang–Mills — Fail
Navier–Stokes — Fail
BSD — Fail
Hodge — Fail
Topo — Fail
|
Artificial Analysis Intelligence Index 36, September 2026. Fails all seven tracks. |
Discussion
Outlook
Every released model fails every track. Benchmark scores tend to rise quickly once a capability becomes a training target, so pass rates here will probably rise too. OpenAI’s Navier–Stokes result shows that a model outside this leaderboard can already resolve one of the actual problems. Passing a track shows strong closed-ended mathematical reasoning under strict verification. It does not show autonomous research ability or general intelligence. The benchmark covers structured proof-style problems, not open-ended research.
Purpose
The benchmark gives researchers, labs, and policymakers one pass-or-fail reference point for mathematical reasoning. That supports concrete discussion of capability trends, risks, and governance. Tracking these results alongside progress on the actual problems shows where current systems succeed, where they fail, and what has changed.
Contact & external results
Send results from your own runs, corrections, or updated numbers for a model.