MillenniumPrizeProblemBench

MillenniumPrizeProblemBench
How frontier models do on the seven Millennium Prize Problems.

Each track is a synthetic task suite modeled on one Millennium Prize Problem. Tasks cover proof search, conjecture generation, formal verification, and research-level reasoning. Results are pass or fail per track.

13 models — Claude, GPT, Gemini, Grok, Qwen, DeepSeek…
7 tracks — one for each Millennium Prize Problem.
Pass / Fail only — no partial credit on any track.

Status: no released model passes any track. An unreleased OpenAI internal model has resolved the actual Navier–Stokes problem, outside this benchmark.

Tracks

Click a track for the problem statement and what the benchmark tests.

P vs NP
Expand
Problem: Decide whether every problem whose solution can be verified quickly (NP) can also be solved quickly (P), or give a proof that P ≠ NP.
Benchmark: reductions, proof sketches, and complexity arguments. Tasks do not ask the model to resolve P vs NP.
Riemann Hypothesis
Expand
Problem: Show that all nontrivial zeros of the Riemann zeta function lie on the critical line Re(s) = 1/2.
Benchmark: analytic number theory tasks on zero distributions, L-functions, and conjecture generation.
Yang–Mills / Mass Gap
Expand
Problem: Construct a quantum Yang–Mills theory on four-dimensional spacetime and prove the existence of a positive mass gap.
Benchmark: simplified PDE and field-theory tasks on gauge symmetry, energy bounds, and mass-gap arguments.
Navier–Stokes
Expand
Problem: Prove or disprove global existence and smoothness for solutions of the 3D incompressible Navier–Stokes equations with smooth initial data.
Status: OpenAI reported a disproof on September 8, 2026: a finite-time singularity under smooth forcing, with a Lean formalization. OpenAI says it will not claim the prize.
Benchmark: simplified fluid PDE tasks on blow-up, regularity, and existence arguments.
Birch & Swinnerton-Dyer
Expand
Problem: Relate the arithmetic of an elliptic curve (its rank) to the order of vanishing of its L-function at s = 1.
Benchmark: tasks on elliptic curves, rational points, and L-functions.
Hodge Conjecture
Expand
Problem: Determine whether certain cohomology classes on projective algebraic varieties are algebraic cycles.
Benchmark: tasks in cohomology and algebraic geometry patterned on Hodge-theoretic arguments.
Topological Surrogates
Expand
Problem: The Poincaré conjecture on 3-manifolds, already proved by Perelman. This track stands in for open problems in topology.
Benchmark: 3-manifold and homotopy tasks.

Model leaderboard

Pass or fail per model and track. No released model passes any track. The OpenAI internal model row records a result on the actual problem, not a benchmark run.

Model P vs NP Riemann Yang–Mills Navier–Stokes BSD Hodge Topo Summary Notes
OpenAI internal model
OpenAI · Unreleased
1 / 7
N/A N/A N/A Resolved N/A N/A N/A
P vs NP — Not evaluated Riemann — Not evaluated Yang–Mills — Not evaluated Navier–Stokes — Resolved BSD — Not evaluated Hodge — Not evaluated Topo — Not evaluated
Reported by OpenAI on September 8, 2026. An internal model, run as thousands of coordinating agents, proved that 3D incompressible Navier–Stokes flow from smooth rest under smooth forcing can form a finite-time singularity with finite energy. This settles statements C and D of the Clay formulation. The proof has a Lean formalization checked with GPT-6 Astra. The model is unreleased, not on the Artificial Analysis leaderboard, and has not run on this benchmark. OpenAI says it will not claim the prize.
Claude Fable 5.1
Anthropic · Sep 2026
0 / 7
Fail Fail Fail Fail Fail Fail Fail
P vs NP — Fail Riemann — Fail Yang–Mills — Fail Navier–Stokes — Fail BSD — Fail Hodge — Fail Topo — Fail
Artificial Analysis Intelligence Index 53, September 2026. Fails all seven tracks.
GPT-6 Astra
OpenAI · Sep 2026
0 / 7
Fail Fail Fail Fail Fail Fail Fail
P vs NP — Fail Riemann — Fail Yang–Mills — Fail Navier–Stokes — Fail BSD — Fail Hodge — Fail Topo — Fail
Artificial Analysis Intelligence Index 53, September 2026. Fails all seven tracks.
Claude Opus 5
Anthropic · Jul 2026
0 / 7
Fail Fail Fail Fail Fail Fail Fail
P vs NP — Fail Riemann — Fail Yang–Mills — Fail Navier–Stokes — Fail BSD — Fail Hodge — Fail Topo — Fail
Artificial Analysis Intelligence Index 51, September 2026. Fails all seven tracks.
Muse Spark 1.3
Meta · Sep 2026
0 / 7
Fail Fail Fail Fail Fail Fail Fail
P vs NP — Fail Riemann — Fail Yang–Mills — Fail Navier–Stokes — Fail BSD — Fail Hodge — Fail Topo — Fail
Artificial Analysis Intelligence Index 48, September 2026. Fails all seven tracks.
GPT-5.6 Sol
OpenAI · Jul 2026
0 / 7
Fail Fail Fail Fail Fail Fail Fail
P vs NP — Fail Riemann — Fail Yang–Mills — Fail Navier–Stokes — Fail BSD — Fail Hodge — Fail Topo — Fail
Artificial Analysis Intelligence Index 47, September 2026. Fails all seven tracks.
GLM-5.3
Z AI · Aug 2026
0 / 7
Fail Fail Fail Fail Fail Fail Fail
P vs NP — Fail Riemann — Fail Yang–Mills — Fail Navier–Stokes — Fail BSD — Fail Hodge — Fail Topo — Fail
Artificial Analysis Intelligence Index 45, September 2026. Fails all seven tracks.
Grok 4.6
xAI · Aug 2026
0 / 7
Fail Fail Fail Fail Fail Fail Fail
P vs NP — Fail Riemann — Fail Yang–Mills — Fail Navier–Stokes — Fail BSD — Fail Hodge — Fail Topo — Fail
Artificial Analysis Intelligence Index 44, September 2026. Fails all seven tracks.
Kimi K3
Kimi · Jul 2026
0 / 7
Fail Fail Fail Fail Fail Fail Fail
P vs NP — Fail Riemann — Fail Yang–Mills — Fail Navier–Stokes — Fail BSD — Fail Hodge — Fail Topo — Fail
Artificial Analysis Intelligence Index 44, September 2026. Fails all seven tracks.
Gemini 3.8 Flash
Google DeepMind · Sep 2026
0 / 7
Fail Fail Fail Fail Fail Fail Fail
P vs NP — Fail Riemann — Fail Yang–Mills — Fail Navier–Stokes — Fail BSD — Fail Hodge — Fail Topo — Fail
Artificial Analysis Intelligence Index 41, September 2026. Fails all seven tracks.
Qwen3.8 Max
Alibaba · Aug 2026
0 / 7
Fail Fail Fail Fail Fail Fail Fail
P vs NP — Fail Riemann — Fail Yang–Mills — Fail Navier–Stokes — Fail BSD — Fail Hodge — Fail Topo — Fail
Artificial Analysis Intelligence Index 40, September 2026. Fails all seven tracks.
Claude Sonnet 5
Anthropic · Jun 2026
0 / 7
Fail Fail Fail Fail Fail Fail Fail
P vs NP — Fail Riemann — Fail Yang–Mills — Fail Navier–Stokes — Fail BSD — Fail Hodge — Fail Topo — Fail
Artificial Analysis Intelligence Index 38, September 2026. Fails all seven tracks.
DeepSeek V4 Pro
DeepSeek · Aug 2026
0 / 7
Fail Fail Fail Fail Fail Fail Fail
P vs NP — Fail Riemann — Fail Yang–Mills — Fail Navier–Stokes — Fail BSD — Fail Hodge — Fail Topo — Fail
Artificial Analysis Intelligence Index 36, September 2026. Fails all seven tracks.
Methodology: each track is a synthetic task suite that copies the structure of one Millennium Prize Problem: formal proof steps, conjecture search, counterexample discovery, and self-critique. Automated checkers and human experts grade the results. A model passes a track only if it passes consistently. No released model listed here has solved a genuine Millennium Prize Problem. The OpenAI internal model row is included because its Navier–Stokes result is a resolution of the actual Clay problem. Its Resolved mark is not a benchmark pass, and that model has not run on any track.

Discussion

MillenniumPrizeProblemBench tracks how far frontier models are from the reasoning the Millennium Prize Problems demand.

Outlook

Every released model fails every track. Benchmark scores tend to rise quickly once a capability becomes a training target, so pass rates here will probably rise too. OpenAI’s Navier–Stokes result shows that a model outside this leaderboard can already resolve one of the actual problems. Passing a track shows strong closed-ended mathematical reasoning under strict verification. It does not show autonomous research ability or general intelligence. The benchmark covers structured proof-style problems, not open-ended research.

Purpose

The benchmark gives researchers, labs, and policymakers one pass-or-fail reference point for mathematical reasoning. That supports concrete discussion of capability trends, risks, and governance. Tracking these results alongside progress on the actual problems shows where current systems succeed, where they fail, and what has changed.

Contact & external results

Send results from your own runs, corrections, or updated numbers for a model.