For most of its history, mathematics has been a story about scarcity — the scarcity of genius. A Ramanujan intuition. An Erdős conjecture that sits unsolved for eighty years, waiting for the one mind capable of cracking it. Progress depended on rare individuals having rare insights, often after decades of quiet obsession. That model just broke, and it broke fast.
Over the last eighteen months, general-purpose AI models — paired with formal proof-verification systems — have started producing research-level mathematical results not through flashes of inspired intuition, but through brute-force exploration, statistical pattern-matching across the literature, and iterative self-correction. Mathematics is starting to look less like an art practiced by the gifted few and more like an engineering discipline: a function of model scale, inference budget, and search strategy.
The Timeline Compressed a Decade Into Two Years
What might once have unfolded over a generation happened in roughly twenty-four months:
Summer 2025 — AI models solved five of six problems at the International Mathematical Olympiad, a result that stunned mathematicians who hadn’t expected that level of competence so soon.
Late 2025 — DeepMind’s AlphaEvolve, which uses Gemini to write and “evolve” Python programs through genetic algorithms, was put to the test by Terence Tao and collaborators across 67 open problems. It improved on the best-known solution in 23 of them — work that would normally take an expert months, compressed into a day or two per problem.
January 2026 — While probing unrelated territory, AlphaEvolve stumbled onto a previously unnoticed hypercube structure buried in fifty-year-old permutation-group mathematics. Mathematician Geordie Williamson noted that this kind of discovery would once have demanded “a real engineering effort” in coding and tuning — now it took twenty minutes.
February 2026 — The “First Proof” challenge presented AI models with ten research-level problems deliberately excluded from their training data. The models solved more than half. Some called it the moment AI “finished graduate school.”
May 2026 — An unreleased OpenAI internal model cracked the eighty-year-old Erdős planar unit distance problem.
July 2026 — Anthropic’s Claude reportedly disproved the Jacobian conjecture with a compact counterexample, overturning decades of failed proof attempts by human mathematicians.
August 1, 2026 — OpenAI published a 249-page manuscript in which its unreleased “Astra” model resolved ten long-standing open problems — spanning sphere packing, error-correcting codes, network structure, quantum game theory, and non-sofic groups — each one backed by a machine-checkable Lean 4 proof certificate.
September 8, 2026 — OpenAI announced the resolution of the Navier-Stokes existence and smoothness problem, one of the seven Clay Mathematics Institute Millennium Prize Problems, each carrying a $1 million reward. An internal model “significantly more capable than GPT-6 Astra” coordinated roughly 10,000 autonomous agents, which produced a proof in 88 hours that a smooth fluid at rest, under a smooth applied force with finite energy throughout, can develop a singularity in finite time — the fluid effectively “blowing up” in a spiraling, spaghetti-like vortex. Another 17 hours went into formalizing the result in Lean. The whole exercise cost several million dollars in compute, according to OpenAI’s Sébastien Bubeck.
Why Compute Now Substitutes for Genius
Three structural shifts explain the transition from chance-and-genius to models-and-compute.
Formalization became infrastructure. Proof-verification languages like Lean and Coq convert mathematics from persuasive prose into machine-checkable code. This is the load-bearing innovation. An AI-generated proof without a Lean certificate is just an assertion; with one, it carries the same rigor as a peer-reviewed paper. Verification stopped being a matter of trusted human judgment and became a mechanical, infinitely scalable process. The Navier-Stokes proof made this literal: OpenAI didn’t just publish a write-up, it published a Lean formalization alongside it, and a second AI model handled verification.
Statistical exploration replaced isolated insight. Where a single mathematician once worked one problem at a time, researchers can now, in Tao’s words, “solve thousands of problems at once and start doing statistical studies.” Discovery becomes a search-and-optimization exercise over a landscape of known results, rather than a singular creative leap. OpenAI’s Navier-Stokes push embodied this at industrial scale: nearly 100 agents spent 50 hours producing a related Euler-equation disproof, then roughly 10,000 agents attacked Navier-Stokes itself, exchanging almost 5 million messages with each other along the way. That’s not a eureka moment — it’s a coordinated swarm doing distributed search.
Cost replaced talent scarcity as the binding constraint. OpenAI estimated that generating Astra’s ten August solutions cost roughly $2,000 in API tokens; the Navier-Stokes run, using a far larger agent swarm, ran into the millions. Either way, solving a world-class open conjecture became a compute-purchasing decision rather than a talent-acquisition one. That’s the clearest signal of the engineering reframe: mathematical progress now scales with budget and inference time the way chip fabrication scales with capital, not with the unpredictable arrival of a singular mind.
Recombination, Not Yet Revolution
It’s worth being precise about what’s actually new here — because most of it isn’t spontaneous originality. Mathematicians reviewing these results describe the models as excelling at finding buried references, connecting disparate subfields, and pushing known techniques further, rather than inventing conceptually new mathematics.
The Navier-Stokes result is the sharpest illustration yet. The real intellectual breakthrough belongs to two human mathematicians, Diego Córdoba and his former doctoral student Luis Martínez-Zoroa, who since Martínez-Zoroa’s 2021 dissertation had pioneered a purely analytic technique — building an “infinite cascade” of non-singular layered solutions that combine into one singular solution — without using computers at all. Their approach got the field to the threshold; both competing AI efforts leaned on it to cross the finish line. Charles Fefferman, who wrote the Clay Institute’s official problem description, called Córdoba and Martínez-Zoroa “the heroes of the story.” Tristan Buckmaster, whose own AI-assisted effort with Anthropic’s Levent Alpöge was racing OpenAI on a related result, went further, saying Martínez-Zoroa “deserves a Fields Medal.”
That parallel effort also exposed the messier side of the new pipeline. Buckmaster and Alpöge had spent nearly a year working with LLMs before reaching a verified proof for the Euler equations in August — and Buckmaster candidly admitted one of their own papers “can only be described as AI slop.” OpenAI, in turn, acknowledged its Navier-Stokes push was triggered by rumors that Buckmaster and Alpöge were close to a Millennium Prize result, and the two camps have since disputed credit and the sequence of events. Oxford’s James Maynard, reviewing OpenAI’s earlier Astra results, put the caveat plainly: impressive, but not yet clear evidence of “profound new mathematical ideas” of Fields Medal caliber.
The honest framing: these systems are extraordinary at combinatorial search and cross-referencing the accumulated mathematical literature at superhuman speed. They are not yet demonstrating the multi-year strategic vision that produces genuinely new conceptual frameworks — they’re finishing proofs that human mathematicians largely set up.
The Field Is Straining Under the New Model
Mathematics has historically been a low-cost, open, individually-driven discipline. The compute-driven model is straining that culture in three specific ways:
- Funding pressure is building as universities question grants for problems an AI can now solve cheaply, threatening the pipeline that trains the next generation of mathematicians.
- Access is becoming unequal, with frontier capability locked inside proprietary models controlled by a handful of companies — a sharp break from math’s historically open-source tooling.
- Pedagogy is destabilizing, because the graduate-level training problems that build research intuition are exactly the ones AI now solves fastest, making homework unusable as a form of assessment.
In June 2026, over 3,400 mathematicians signed the Leiden Declaration, urging against AI-company hype that could convince funders human mathematicians are less necessary than they actually are. The Navier-Stokes credit dispute in September only sharpened that anxiety, turning an abstract worry about attribution into a very public, very messy race between two AI labs.
Where the Limits Still Sit
Current systems still cap out around three to four pages of proof (Google’s internal models may soon reach ten), struggle with the long-horizon strategic planning that novel conceptual breakthroughs require, and generate what one mathematician bluntly called an “ocean of slop” alongside genuine advances — requiring careful human or automated verification to separate the two. Even the Navier-Stokes proof, despite its Lean certification, still needs the community to confirm that what was formalized actually matches the Millennium Prize’s precise problem statement — a step Quanta noted remains the crucial piece of human judgment left in the loop.
Tao’s own metaphor captures the moment best: AI today is like a jumping robot that can parkour up small walls of unresolved lemmas, but is “nowhere near the Mount Everests of math” that demand sustained, multi-year strategic vision. Navier-Stokes suggests those walls just got a lot taller — but the climb still started from a ledge human mathematicians built. Whether raw compute scaling eventually closes the remaining gap, or whether it requires a genuinely different architecture, is the question that will define mathematics’ next chapter — and it’s one the field is now confronting not in decades, but in quarters.

