← All articles

OpenAI released almost 400 AI-generated mathematical results: what is known about checking them and a rule to adopt when AI hands you a lot at once

Topics: Science, AI

A stream of sheets with geometric shapes pours out of a glowing cube on the left; on the right, a few desks with lamps hold small stacks, and only some sheets carry green round marks

Short answer: according to The Verge, in October 2026 OpenAI released almost 400 AI-generated mathematical results in more than 700 manuscripts. More than three dozen mathematicians the outlet spoke with described what happened with words such as "staggering" and "unprecedented," and researchers, the outlet writes, agreed that simply understanding what was released could take years. The main results are formalized, that is, written in the Lean language for computer checking, in 300 of 719 manuscripts, about 42%, and OpenAI acknowledges that the results are at different stages of verification. This is not the number of confirmed results: one still has to check that the formalization proves what is claimed. By October 8, the company had removed three papers because of a sign error. Opinions on quality differ: some mathematicians speak of work of a very high level, others complain of hard-to-read texts. For any work with AI, a rule follows: if it hands you many results at once, checking becomes the main bottleneck. Below are the facts, the opinions and three steps.

What was published

All the facts below come from The Verge's article; I did not open OpenAI's repository.

  • Volume. Almost 400 results in more than 700 manuscripts, covering combinatorics, geometry, number theory, theoretical computer science, algebra, topology, probability and statistical mechanics, and mathematical physics. The table of contents with abstracts runs to about 40 pages, and for several mathematicians reading it alone took about an hour.
  • How OpenAI got them. The model attempted more than 4,000 problems, and a typical result took about three hours of "thinking" compute in ChatGPT Pro. The company did not disclose the model, the prompts or the full list of problems.
  • How they are checked. Some manuscripts come with formalizations in Lean, a language in which a computer checks a proof. But, according to mathematicians, one still has to make sure the formalization proves exactly what the text claims, and the quality is uneven and the statements do not always match what is claimed.
  • What is missing. According to OpenAI, the main results are formalized in 300 of 719 manuscripts, about 42%. For the remaining 419 manuscripts, the article does not name formalizations of the main results and does not say how they were checked otherwise. Formalization, in turn, does not remove the need to check that it proves what is claimed.
  • Corrections. By October 8, OpenAI's correction log listed revisions to more than a dozen manuscripts and the removal of three because of a sign error which, according to the company, invalidates the argument.

What mathematicians say

  • On the quality of the work. A number of researchers believe the level is very high despite weak presentation. In the view of some, before the AI era many of the results would have deserved publication in leading journals, and a handful could make their author a contender for the Fields Medal, one of the highest awards in mathematics. Stanford mathematician Jared Duker Lichtman named "tens" of such results, including progress toward the Riemann hypothesis, a special case of the Hodge conjecture and a solution to the four-dimensional Kakeya conjecture. This is his opinion, not confirmed by verification.
  • On the texts. Brendan Hassett of Brown said that the write-up of the problem he knows best made little sense after a quick read. He said this before three papers were retracted. Experts criticized OpenAI's earlier mathematical publications, in particular for weak attribution of sources.
  • On checking. Kevin Buzzard of Imperial College London identified about six theorems in his field and said he has to either read possibly incorrect texts or wait for others to check them or formalize them.
  • On consequences. Some researchers said that colleagues' plans and grant proposals had been devalued: Scott Armstrong knows of a group whose research program has practically been "wiped out," and Tristan Buckmaster heard of three people in that situation.
  • On significance. Constantin Kogler of the Institute for Advanced Study called it the most important moment in the history of mathematics, and Yang-Hui He compared this year's events to the publication of Euclid's Elements. Most mathematicians were less categorical.

OpenAI, according to the article, disclosed more than it did on past occasions and keeps a correction log. An advisory group of mathematicians recommended releasing papers that a human understands, formalizing proofs and disclosing models and prompts; OpenAI followed some of this advice but did not name the model and the prompts. The company promised to fund workshops and conferences without giving details and said it would improve the quality of its papers.

I do not assess whether the results themselves are correct: that takes verification by specialists.

Three rules for those who receive many results from AI at once

These are my conclusions from the case described, not the mathematicians' words.

  1. Decide in advance what counts as checked. In this story there are several different states: "generated by AI," "formalized" (written for computer checking), "the formalization has been checked against the claim it is meant to prove" and "correctness confirmed by a specialist." Something read and understood by a person is not the same as confirmed. The status should be written next to each result.
    • Who does it: whoever accepts work from AI.
    • How to check: the table of results has a "verification status" column with one of these values, and there are no empty cells.
  2. Do not publish what cannot be checked in a reasonable time. If checking one result takes hours and there are hundreds of results, the queue for checking can grow if new results arrive faster than they are checked.
    • Who does it: the manager.
    • How to check: estimate the checking time from two or three results and multiply by their number before the work starts. If the total is more than the time you have, take a smaller batch.
  3. Keep a correction log. OpenAI does this, and the log has already shown the retraction of three papers.
    • Who does it: whoever publishes the results.
    • How to check: every correction is recorded with a date and a reason, the previous version is kept, and a reader can find the log in a minute.

A calculation with illustrative numbers

Suppose a person needs 2 hours on average to check one manuscript: this is an illustrative figure, the source does not give it. Then for 719 manuscripts that comes to 1,438 hours, about 36 working weeks of 40 hours for one specialist. For your business the example is the same: if AI has prepared a hundred contracts, reports or calculations and checking each takes half an hour, checking will take 50 hours. Substitute your own numbers and compare with the time you have.

When this does not concern you

If you ask AI to write a couple of letters or one table, you can check the result right away. The rules above are needed when there are many results and an error in one is hard to notice among the rest.

Summary

In October 2026, OpenAI released almost 400 AI-generated mathematical results in more than 700 manuscripts. The main results are formalized in 300 of 719 manuscripts, about 42%; three papers have been removed, and mathematicians' opinions range from "a very high level" to "hard to read." I do not assess whether the results themselves are correct. For those who work with AI, the main point of this case is that checking becomes the bottleneck, so the verification status should be written next to each result, the batch should be chosen by checking time, and a correction log should be kept.

I work on AI agents and automation. If you want to look at my projects or discuss your own task, visit my portfolio.

Sources