On October 7, OpenAI’s repository history recorded the withdrawal of three mathematics manuscripts after a sign error invalidated a key argument and a construction used by two dependent papers. The entry also listed revisions, citation updates, and additional Lean formalizations, a day after OpenAI released its mathematics-results repository and the Association for Human Mathematics (AHM) criticized the scale of that release.

OpenAI withdrew three math manuscripts

OpenAI’s repository history says the sign error affected a stabilization-trace cancellation argument and a construction used by two dependent papers. OpenAI withdrew three manuscripts:

  • “Algebraicity of Weil classes on split abelian eightfolds”
  • “Algebraicity of Kuga–Satake Correspondences for K3 Surfaces”
  • “The rational Hodge conjecture for products of K3 surfaces”

The entry identifies the error and the connected work behind these withdrawals. They concern those manuscripts; the entry does not assign the same outcome to every result in the repository.

This is a separate development from the earlier OpenAI Navier–Stokes claim. The October 6 release covered a broader collection of mathematical results.

What the AHM objected to

OpenAI withdraws three math manuscripts after sign error

The AHM criticized the scale and publication approach, saying OpenAI released more than 700 files at once and urging mathematicians to discontinue their work with the company. Its objection centered on the release as a bulk publication, not on a count of formally checked results. The AHM statement sets out the organization’s position.

OpenAI’s stated publication approach

OpenAI’s October 6 announcement described a repository of results generated with an internal model, with paper-review and citation protocols. The company said it had consulted the independent Institute for Advanced Study Advisory Group on Mathematics and Artificial Intelligence and was incorporating public recommendations into its publication practices.

OpenAI also said the release included reasoning summaries, compute estimates, and information about attempted problems. It reported an average of about three hours of ChatGPT Pro-equivalent reasoning compute per result. That is OpenAI’s estimate for its release, not a measure of review or acceptance by mathematicians.

What the Lean count measures

Lean is a proof assistant: mathematicians can express a proof in a precise formal language that software can check against encoded definitions and rules. OpenAI’s October 7 history lists 300 of 719 top-line results as formalized in Lean—about 42% of that dated repository count.

That number tracks which results had Lean formalizations in the October 7 entry. It is not a count of results accepted by the mathematical community. Formal checking and mathematical review answer different questions: Lean checks a proof as encoded, while mathematicians assess the result and its role in further work.

The dispute is not simply about whether software can check formal statements. It is also about how researchers can examine and use a large release, and whether its presentation supports mathematical understanding.

Other changes recorded on October 7

Alongside the three withdrawals, OpenAI’s repository history recorded revisions to 14 manuscripts, citation updates in 13 additional manuscripts, and six additional formalizations. These are separate categories of changes in the same dated entry.