Skip to content
AI Primer
breaking

OpenAI publishes 722 math manuscripts from an internal model

OpenAI published 722 manuscripts covering 372 mathematical result families from about 4,000 problems. Independent checks found that some headline claims were partial results rather than complete proofs.

6 min read
OpenAI publishes 722 math manuscripts from an internal model
OpenAI publishes 722 math manuscripts from an internal model

TL;DR

  • OpenAI published 722 manuscripts grouped into 372 mathematical result families, drawn from an evaluation of roughly 4,000 problems, through OpenAI's announcement.
  • The average selected result used compute equivalent to about three hours of ChatGPT Pro thinking, according to rohanpaul_ai's release summary.
  • Lean covers a subset of the drop: WesRoth's post says many proofs were translated, while CellCog's inventory counts 162 papers with a formalized main result.
  • Several headline claims are narrower than their social-media packaging: imjustnewatai's breakdown separates a zero-free strip from the Riemann hypothesis and a CM case from the full Hodge conjecture.

The official release includes a full repository, reasoning summaries, and Lean artifacts. The repository README also says some unformalized results could have issues. A repo-based scope check found several formal statements narrower than their manuscript headlines.

The release

OpenAI published manuscripts produced by an unreleased internal model in the public openai/math repository. The repository contains preprints, Lean code, reasoning traces, a catalogue, and an overview of the result families.

The catalogue is organized around families rather than isolated PDFs. OpenAI says a family can contain a principal result, companion arguments, consequences, or alternative proofs.

  • 722 manuscripts
  • 372 result families
  • 17 mathematical fields, according to CellCog's repository inventory
  • 10 abridged reasoning summaries
  • Supporting Lean formalizations for many, but not all, results

The release also includes revision and citation protocols. OpenAI says it consulted the independent Advisory Group on Mathematics and Artificial Intelligence at the Institute for Advanced Study before publishing the collection.

The evaluation funnel

The model was given approximately 4,000 research problems after OpenAI's existing mathematical evaluations saturated. The repository says the vast majority of results came from one standard procedure, then were grouped into families and manuscripts after OpenAI applied a significance threshold.

The compute figure is an average, not a per-paper budget: OpenAI reports roughly three hours of ChatGPT Pro thinking compute for each result. The repository names work on a Riemann zeta zero-free region and the Hodge conjecture for CM abelian varieties as exceptions to the fixed procedure.

The README also says the writeup for the Re(s) > 11/12 zero-free region was human edited for readability. That detail makes the release a mixed corpus of model outputs, formal artifacts, and selected editorial intervention, rather than a uniform benchmark dump.

Lean formalizations

OpenAI describes Lean as a way to check mathematical proofs by computer and says it will add more formalizations over time.

A repository-based count found Lean formalizations for the main result of 162 papers, leaving 560 manuscripts without a formalized main result in that snapshot. OpenAI's README warns that some of the unformalized results could have issues.

Lean acceptance applies to the formal theorem and proof supplied to the checker. It does not establish that the formal theorem matches every claim in the prose, that the result is novel, or that a manuscript's informal extensions follow from the checked core, according to an independent scope analysis.

The distinction appears in the repository metadata. willdepue's audit update notes that not every result in the public list is Lean verified, even as the release presents the manuscripts together in one catalogue.

Headline scopes

The first pass through the papers produces several concrete scope corrections.

  • Quasi-Riemann hypothesis: The claim is a zero-free half-plane for the zeta function and other L-functions when Re(s) > 7/8. That is a substantial fixed strip, but the classical Riemann hypothesis concerns the critical line at Re(s) = 1/2. The preprint does not claim the full Riemann hypothesis.
  • Hodge conjecture: The manuscript concerns the rational Hodge conjecture for CM abelian varieties, a structured special case of a much broader problem. imjustnewatai's Hodge explainer explicitly describes it as a family-level result rather than a solution of the full conjecture.
  • Matrix multiplication: The reported bound lowers the exponent to 2.25, from the previously cited bound near 2.371. rohanpaul_ai's matrix-multiplication summary adds the engineering caveat: the result is a theoretical asymptotic bound and does not provide a faster routine for real GPU workloads.
  • Top-ranked open problems: A quick catalogue check by LechMazur's ranking check marked the Riemann, Hodge, and Birch and Swinnerton-Dyer entries as partial, while marking other entries as full claims. The list therefore mixes complete claims, partial progress, and different levels of formal support under the same release.

The same problem appears in the repository's selected headline results. A second scope check found a Fourier-transform formalization whose Lean statement was narrower than the manuscript claim, while the matrix result was explicitly limited to arithmetic complexity and asymptotic behavior.

The advisory protocol

The advisory group says it operates independently of AI companies, accepts no payment, and has no decision-making power over them. Its stated purpose is to advise on responsible presentation and release of mathematical results.

OpenAI says the group informed its release practices and that it is exploring community-hosted alternatives to the company-controlled GitHub repository. The model that produced the papers remains unreleased; OpenAI says it is working toward a responsible release.

The company also says it will fund workshops, conferences, and special programs focused on understanding major AI-produced mathematical results. The repository keeps the papers public while the model and the broader generation process remain partly closed.

Verification load

The first bottleneck was producing candidate proofs. The next one is reading and checking them. [src:3.1|rohanpaul_ai's verification note] describes formal verification and expert review as the scarce inputs after the proof-generation step.

A repo-based review of eight headline results found that none was described as refereed, even when the repository listed a Lean entry. The distinction matters most where a formal statement covers only a special case or a weaker asymptotic claim.

The scale is visible in the community response. LechMazur's audit warning pointed to faulty proofs elsewhere on arXiv and said the OpenAI list would require actual audits before its status notes could be trusted. Meanwhile, willdepue's audit update reported that the first pass had not yet found a profound error in the OpenAI results, while still expecting some claims to fail scrutiny.

Versioned manuscripts

The repository treats the release as a versioned research corpus. OpenAI says corrections and revisions will be recorded as new versions, with earlier versions remaining accessible, and each manuscript directory includes its own citation block in the repository.

That makes future changes attributable at the paper level: a corrected proof can remain linked to the original claim, rather than silently replacing it. The public collection is therefore both a set of mathematical manuscripts and a record of how their formal status changes over time.

Further reading

Discussion across the web

Where this story is being discussed, in original context.

On X· 8 threads
TL;DR1 post
The release1 post
The evaluation funnel1 post
Lean formalizations3 posts
Headline scopes3 posts
The advisory protocol2 posts
Verification load4 posts
Versioned manuscripts1 post
Share on X