OpenAI releases new math results from an internal model, with Lean proofs

OpenAI releases new math results from an internal model, with Lean proofs

OpenAI has published a post announcing a broad range of new mathematical results produced by an internal frontier model. The model is not named. The post does not describe any individual result, theorem or problem; it is about how the results are being shared.

OpenAI says it has been consulting the independent Advisory Group on Mathematics and Artificial Intelligence at the Institute for Advanced Study to develop best practices for sharing results with the math community, and that it drew on the group's advice and public recommendations in shaping this release.

The results are published in a GitHub repository, with protocols for paper revisions and citations. OpenAI says it is still exploring community-hosted alternatives that meet the committee's guidelines. For future releases it commits to improving the papers through better citations, mathematical exposition and presentation.

The repository includes Lean formalizations of many of the proofs. Lean is a programming language that lets a computer check mathematical proofs. The formalizations cover many of the proofs, not all, and OpenAI says it will add more as it obtains them.

For transparency, the repository also carries details of how the results were obtained: 10 summaries of the model's reasoning, estimates of compute spent in terms of Pro usage on ChatGPT, and statistics on the number of attempted problems. The average result used the equivalent compute of roughly three hours of ChatGPT Pro thinking.

OpenAI also says it will fund a series of workshops, conferences and special programs around understanding major results produced by AI, with more details to come soon. Finally, it says it wants to put state-of-the-art capabilities directly in scientists' hands and is working to responsibly release the model that produced these results. It adds that it will keep acting on community feedback and update its standards for future disclosures of major scientific advances.

Key facts

  • OpenAI is releasing a broad range of new mathematical results from an internal frontier model, which the post does not name.
  • Results go into a GitHub repository with protocols for paper revisions and citations, following advice from the independent Advisory Group on Mathematics and Artificial Intelligence at the Institute for Advanced Study.
  • Many of the proofs come with Lean formalizations that a computer can check; more will be added as they are obtained.
  • The repository includes 10 summaries of the model's reasoning, compute estimates in ChatGPT Pro terms and statistics on attempted problems; the average result used about three hours of ChatGPT Pro thinking.
  • OpenAI says it will fund workshops, conferences and special programs on understanding AI-produced results, and is working to release the model responsibly.

Why it matters

The post is a disclosure protocol as much as an announcement. A lab is pairing claimed mathematical results with machine-checkable Lean proofs, reasoning summaries and compute figures, and says it follows advice from an independent advisory group at the Institute for Advanced Study. Proofs that a computer can verify are a stronger basis for trust than a lab's own say-so. The post also signals that the model behind the results may eventually be released, though it gives no date.

Who it affects

Mathematicians and the wider math community are the main audience: the results, papers and formalizations are meant for them, and OpenAI says it is trying to improve how it shares work with them. Researchers in other sciences are addressed too, since OpenAI says it wants to empower scientists with state-of-the-art capabilities. Groups thinking about how AI-produced results should be cited, revised and evaluated are affected as well, because the post lays out one lab's approach.

How to use it

The results, the Lean formalizations, the reasoning summaries, the compute estimates and the problem statistics are all in the GitHub repository. The post does not give the repository URL in its text. Papers there come with protocols for revisions and citations. Lean formalizations let a reader check proofs by machine rather than by hand, and OpenAI says it will add more as it gets them. The model itself is not available yet; the post says only that OpenAI is working to release it responsibly.

How solid is it

This is a first-party announcement. The source does not say whether the results have been peer reviewed or independently verified. No individual results are described, so their significance cannot be judged from the post. It does not name the model, give the number of attempted problems or the number of results, or say whether the average compute figure of roughly three hours of ChatGPT Pro thinking is a mean or a median. It also does not say the Advisory Group endorsed or reviewed these specific results; it says OpenAI drew on the group's advice and public recommendations. Lean formalizations cover many of the proofs, not all.

Risks and caveats

No release date or timescale is given for the model's public release, the workshops or the additional formalizations; these are stated intentions. Proofs without a Lean formalization still need checking by other means. OpenAI itself says that for future releases it wants to improve paper quality through citations, exposition and presentation, which implies the current papers are a work in progress. Readers should wait for outside mathematicians to examine the repository before treating the results as established.

“We’re releasing a broad range of new mathematical results produced by an internal frontier model.”

— OpenAI, Sharing AI progress in mathematics