A mathematician imagines math progress stalling despite superhuman AI

A mathematician imagines math progress stalling despite superhuman AI

The post is written on the way back from a summit on the future of mathematics held at OpenAI. Sebastian Bubeck had asked the author to speak about a future to avoid, one where humans are mathematically disempowered, and Jacob Tsimerman had advised the speakers to prioritize detail over correctness. The workshop's premise, taken as a given rather than debated, was that AI will become robustly superhuman at mathematics. The author sets out to imagine a story in which, despite that, mathematical progress stalls. This is explicitly framed as not a prediction: the author calls themselves optimistic by nature and expects the field to adapt, but wants to trace where certain existing trends could lead if they continue.

The scenario runs year by year. In 2026, mathematical output is exploding: the number of combinatorics papers posted per week to arXiv has been climbing sharply since late 2021, with other areas showing a similar but less dramatic rise. At the same time, activity on MathOverflow, both questions and answers, has been declining for a while, a decline that has accelerated since the start of 2025 and that the author attributes largely to AI; the author found no offsetting increase in answers to older questions. Early signs of a stranger problem also appear: three groups independently produced very similar proofs of Feige's 1/e conjecture almost simultaneously, two of them disclosing that AI had found the result. Separately, after the account @alpoge posted Fable's counterexample to the Jacobian conjecture in dimension 3 or higher, an internal OpenAI model independently replicated it, and Anthropic likewise replicated several other recent results OpenAI had announced. The pattern the author draws out is that models and the people using them keep converging on the same problems, so a large share of the resulting work, in humans and in AI alike, adds essentially nothing beyond the cost of the tokens spent producing it.

By 2027 in the scenario, the mathematics profession still rewards people for producing papers and resolving conjectures, even though those outputs are now mispriced by the glut of AI-assisted work. If that continues, the author argues, the dominant strategy for career success becomes what they call playing the slot machine for conjectures: a researcher can hand a tool such as Codex the job of picking conjectures, resolving them and checking the work, and produce many short papers a day this way, whether or not they care about correctness, and people are already doing both versions of this. The value added is questioned directly: not expertise, since none is developed, and often not even careful reading, since the author suggests no one, not even the nominal author of a given paper, is reading much of this output. Mathematicians grow disconnected from the mathematics itself, and even human verification becomes less valuable as the models improve. A further chilling effect appears: since models can now reconstruct a paper from a few key ideas, and some autonomous AI results already show a last mile pattern of finishing off problems shortly after others' groundwork, talking about work in progress becomes risky under a system that still rewards priority with prestige and jobs. The author reports being told by multiple colleagues that they are now unwilling to discuss work in progress for exactly this reason.

2028 brings a genuine bright spot in the scenario: autoformalization, translating mathematical statements into machine-checkable formal language, becomes cheap and effective, and many gaps or errors in the existing literature get found and fixed this way. The author is skeptical of the common claim that human judgment remains necessary to confirm statements and definitions are formalized correctly, expecting models to handle that too, but flags a real and separate problem already appearing: cases where a formalization subtly differs from the English text it is meant to formalize, in ways not obvious to a reader. That, too, deepens the disconnection between mathematicians and the mathematics they nominally produce, and the author notes that writing plain-language explanations of formal results alongside them helps but is costly and slow.

From 2029 onward, the scenario has the profession still incentivizing paper production even as models come to fulfill essentially every function human mathematicians perform today: building theory, conjecturing, resolving conjectures, iterating. Human mathematicians are left doing what the author calls lab science with agents, directing compute toward whatever questions interest them. This raises open questions the author does not resolve: who actually engages with the resulting work, how the next generation of mathematicians gets trained, and whether institutions that fail to adapt can keep producing high-quality mathematicians at all. The author suspects existing incentive structures will increasingly reward people who do not engage deeply with the mathematics, or even care about it, and doubts that a sustainable mathematical practice results from that path.

The author frames these problems as symptoms of flaws already present in the profession's institutions, which AI-driven stress will expose rather than create, and suggests this could be an opportunity to fix them: the underlying values, such as producing high-quality science and human understanding, persist, but the mechanisms currently used to reward them are not robust to highly capable AI. The post closes on an optimistic note distinct from the scenario itself: the author expects mathematics to survive and flourish and expects the field to adapt, while acknowledging these concerns may look narrow next to the broader social upheaval highly capable models are likely to cause well beyond mathematics.

Key facts

  • The post follows a summit on the future of mathematics hosted at OpenAI, where Sebastian Bubeck asked the author to speak about a future in which humans are mathematically disempowered.
  • Three groups independently produced very similar proofs of Feige's 1/e conjecture almost simultaneously, two disclosing AI had found the result; separately, an internal OpenAI model and Anthropic each independently replicated other recent AI math results, including a Jacobian conjecture counterexample in dimension 3 or higher first posted by @alpoge.
  • MathOverflow questions and answers have been declining, with the drop accelerating since the start of 2025, which the author attributes largely to AI.
  • The scenario projects that if current incentives (rewarding papers and resolved conjectures) continue unchanged, the dominant career strategy becomes using tools like Codex to pick, resolve and check conjectures automatically, producing many short papers a day.
  • By 2029 and beyond in the scenario, AI models fulfill essentially every function human mathematicians perform today, leaving humans doing what the author calls lab science with agents.

Why it matters

The piece is a direct response to a live premise inside AI labs, that AI will soon be robustly superhuman at mathematics, and asks a question that premise does not answer on its own: whether the institutions around mathematics, built to reward papers and resolved conjectures, can survive that shift intact. The author's answer is that survival is not automatic. Even in a world where AI genuinely gets better at math than people, the argument goes, the profession's current reward structures could push it toward a glut of duplicative, mispriced output and away from the deep engagement that made the field worth doing in the first place.

Who it affects

Working mathematicians and the institutions that structure their careers: journals and preprint servers like arXiv, community venues like MathOverflow, hiring and tenure committees that reward papers and resolved conjectures, and funding bodies that rely on those same signals. It also affects the training pipeline for future mathematicians, since the scenario questions whether current institutions, unchanged, can keep producing high-quality mathematicians once models can do most of the technical work themselves.

How to use it

The essay is not a tool or a product, but it hands readers a concrete diagnostic: watch whether reward structures for research keep tracking paper counts and resolved conjectures once AI can generate both cheaply, and watch whether people become reluctant to discuss work in progress, which the author already reports hearing from colleagues. Institutions or individuals deciding how to weigh AI-assisted mathematical work can use the year-by-year scenario as a checklist of early warning signs, several of which, per the author, are already visible: MathOverflow's accelerating decline, near-simultaneous AI-assisted proofs of the same conjecture from separate groups, and formalizations that quietly diverge from the English statements they are meant to match.

How solid is it

The author is explicit that this is not a prediction: it is a deliberately speculative, non-forecasting exercise in tracing where existing trends might lead, written by someone who describes themselves as optimistic about mathematics adapting. The scenario's individual anchor points, the Feige 1/e conjecture being proved near-simultaneously by three groups, an OpenAI model and Anthropic independently replicating other AI math results, and MathOverflow's decline, are presented as real, current observations rather than invented illustrations, but the specific figures behind the arXiv and MathOverflow trends are not given in the text, and the imagined 2026 through 2029-plus timeline itself is a thought experiment, not a documented sequence of events.

Risks and caveats

The scenario the author sketches includes a rising volume of duplicative work whose marginal value is close to the token cost of producing it; incentive structures that reward paper counts even after those outputs become mispriced; mathematicians who develop no real expertise because no one, including the nominal author, closely reads much of the resulting work; a chilling effect in which researchers become unwilling to discuss work in progress for fear a model will finish it off first; and formalizations that can subtly diverge from the English statements they claim to represent, in ways not obvious to readers. The author frames all of this as an exposure of flaws already present in the profession's institutions rather than something AI invents from scratch, and closes by doubting that a mathematical practice built on these dynamics is sustainable in the long run, while still expecting the field as a whole to adapt and survive.

“I think if we do, the dominant strategy for career success (at least in the medium term) is playing the slot machine for conjectures.”

— the author