OpenAI's Astra claims 10 math breakthroughs, sparking a crisis in mathematics

OpenAI's Astra claims 10 math breakthroughs, sparking a crisis in mathematics

A few weeks before this interview, OpenAI published a blog post, backed by supporting papers that The Verge's London-based AI reporter Robert Hart estimates ran several hundred pages, under the title "10 Advances in Mathematics and Theoretical Computer Science." It credited OpenAI's newest model, Astra, with producing solutions, in some capacity, across an array of disciplines, including quantum game theory and sphere packing in dimensions higher than three, plus several other areas Hart does not name in the interview. The release followed an earlier result in May, when an unnamed internal OpenAI model disproved the unit distance conjecture, an 80-year-old problem. Hart suspects that earlier model was Astra too, but says OpenAI would not confirm this when asked directly.

On The Verge's Decoder podcast, Hart described the reaction from the math community as "a bit of a bombshell." These were not obscure curiosities: mathematicians Hart interviewed said the ten problems were ones the field genuinely cared about and had failed to crack on its own. Several told Hart that solving even one of the ten would typically be enough to set a person up for an academic career, and that solving all ten in a single release would have seemed unbelievable coming from a human mathematician.

The response was not uniformly positive. OpenAI's post initially framed the ten problems as having seen no progress in the past decade, but one of its own accompanying papers states plainly that it builds on progress made by two other researchers on that specific problem. That framing was later revised, quietly and without any announcement. A few researchers Hart interviewed pushed back on the missing credit for work OpenAI had leaned on heavily; one researcher who was actually named in OpenAI's own papers told Hart they felt ambivalent about the whole episode.

Hart's own assessment, after reading the papers, is that this does not amount to plagiarism. One researcher Hart spoke with estimated that only about 50 people worldwide would ever read the material that closely, but Hart says the papers plainly do not attempt to pass off other people's work as original. Hart's verdict is that OpenAI's post was simply a poorly written press release, and Hart adds, from a decade-plus of covering science, that press releases routinely oversell a finding's novelty and importance, calling this another case of that pattern.

Underneath the dispute sits a bigger question the episode keeps returning to: what a run of results like this means for mathematicians' future. Hart frames the disruption as compressed: other fields, like software engineering, have spent roughly five years adjusting to growing AI capability, while math, on Hart's account, went through a comparable shift in just the past six months to a year, after what Hart calls a threshold effect in how well the newest models handle abstract reasoning. The episode poses two questions without answering them: what purpose academic grants and university programs meant to train new mathematicians still serve if frontier models can simply answer the field's open problems, and whether all the attention paid to AI and math is really just a marketing exercise for AI labs.

Hart cautions against treating math as a single skill. The same systems behind these claimed breakthroughs are, Hart says, still bad at basic arithmetic, counting, telling time and identifying the day of the week. Hart's own boyfriend, by Hart's account, keeps noting that the models still think every day is Wednesday, and Hart cites an older, widely repeated failure in which models could not correctly count the number of R's in "strawberry." That specific glitch looks fixed now, and the conversation jokes it may simply be hard-coded rather than a genuine sign of better reasoning, a theory Hart says is one worth buying into. Mathematicians have floated topology as one area where AI is still weak, though Hart says checking that claim sits beyond Hart's own expertise. AI's new competence, on this account, clusters around problems that are self-contained and mechanically checkable rather than tasks needing real-world judgment, which is also why the episode raises, without answering, whether these research skills could transfer to other domains; the interview excerpt itself breaks off with the host asking whether the feat is even repeatable.

Key facts

  • OpenAI's blog post "10 Advances in Mathematics and Theoretical Computer Science," backed by papers Hart estimates ran several hundred pages, credits the company's newest model, Astra, with producing solutions across disciplines including quantum game theory and higher-dimensional sphere packing.
  • The release followed an earlier result in May, when an unnamed internal OpenAI model disproved the 80-year-old unit distance conjecture; Hart suspects it was Astra too but could not get OpenAI to confirm this.
  • OpenAI's post initially claimed no progress on the ten problems in the past decade, but one of its own papers credits two other researchers with earlier progress on that problem, and the claim was later quietly revised.
  • Researchers Hart interviewed were largely impressed, several saying that solving even one of the ten problems would typically be enough for an academic career, while a few objected to the initial lack of credit.
  • The debate has fed a wider argument about mathematicians' future as AI's advanced-math ability grows, even though the same systems remain weak at basic arithmetic, counting and telling time.

Why it matters

A general AI system claiming genuine, expert-level contributions to open problems in one of academia's oldest and most rigorous fields is a different order of claim than passing a benchmark. Hart frames this as the same kind of phase change other fields, like software engineering, went through over roughly five years, except compressed here into six months to a year. If frontier models can answer hard open problems directly, that puts real pressure on the traditional research pipeline, the point the episode itself raises: what academic grants and university programs meant to train new mathematicians are actually for, if the models simply answer the field's outstanding questions.

Who it affects

Mathematicians and math researchers, whose funding, career incentives and daily purpose are the direct subject of the debate; OpenAI, which is staking a research claim in a field it now has to defend against pushback on crediting; and Robert Hart, The Verge's London-based AI reporter, whose original reporting the interview is built around. The researchers who reacted to OpenAI's claims are not named in this account, though one was a person credited within OpenAI's own papers.

How to use it

There is no product or price here; the practical takeaway is where to look and how skeptically. Readers who want to check the substance themselves can go to OpenAI's own "10 Advances" blog post and its accompanying papers, the same material Hart worked from. Hart's own experience is a useful filter here: a lab's press-release framing of its own research, especially a claim like no progress in a decade, is worth checking against the paper itself before repeating it, since that is exactly where this release went wrong.

How solid is it

The mathematics itself is not what is disputed. Researchers Hart interviewed, after reviewing OpenAI's papers, broadly called the results real, difficult achievements; the objection was to OpenAI's framing and crediting, not to whether the problems were actually solved. The one confirmed factual conflict is specific: OpenAI's post said there had been no progress on the ten problems in ten years, while one of its own papers credits two other researchers with earlier progress, and OpenAI quietly revised the claim afterward. Separately, whether the May unit-distance-conjecture result also came from Astra is Hart's own unconfirmed guess; OpenAI did not answer when Hart asked directly.

Risks and caveats

The reaction described rests on a small, unnamed group of researchers Hart interviewed, not a survey of the field. AI's demonstrated competence here clusters around problems that are self-contained and mechanically verifiable; whether that generalizes to messier, real-world domains is a question the episode raises and does not resolve. The same systems remain unreliable at simple tasks, arithmetic, counting, tracking the day of the week; mathematicians have separately flagged topology as an area where AI is still weak, though Hart says verifying that specific claim is beyond Hart's own expertise. The interview excerpt itself ends mid-question, with the host asking whether the feat is repeatable at all, a question this material does not answer.

“If a researcher had done any one of these problems, they'd probably be set for an academic career.”

— researchers interviewed by The Verge's Robert Hart, on the Decoder podcast