OpenAI published 377 new mathematical results on Tuesday, spanning algebra, number theory, theoretical computer science, mathematical logic and topology, in a release that has stunned working mathematicians and reopened a debate about whether AI systems are producing genuine creative work or finishing proofs that human researchers started. The announcement came weeks after the company said it had cracked the Navier-Stokes equation, one of the Clay Mathematics Institute’s seven Millennium Prize Problems, each carrying a $1 million reward, a claim that had already divided the field.
What was actually released
The results were produced by an unreleased internal model. OpenAI said most of the proofs had been formally checked using Lean, a computer language designed to verify that underlying logic is sound, which is the strongest form of verification available for mathematical claims. For 10 of the solutions the company also published summaries showing how the model arrived at the result. The average result took about three hours of computing time.
| Detail | What OpenAI disclosed |
|---|---|
| Results published | 377 |
| Fields covered | algebra, number theory, theoretical CS, logic, topology |
| Verification | Lean proof assistant for most proofs |
| Average compute per result | about 3 hours |
| Model used | unreleased internal system |
| Prompts and traces released | 10 of 377 solutions |
The release follows a pattern OpenAI has established over recent months. Once models exhausted the standard benchmark problems used to test mathematical ability, the company built new benchmarks out of unsolved problems at the cutting edge of active research, and those unsolved problems are what surfaced this week. OpenAI said it drew on public recommendations from an external advisory group in deciding how to publish, and Roberts framed the work as a byproduct of testing rather than the goal itself.
The advisory board tension
A couple of weeks before the release, OpenAI began working with an advisory board of mathematicians hosted by the Institute for Advanced Study in Princeton, New Jersey, independent of any AI company. The board recommended that all prompts given to the AI agents, and the agents’ chains of thought, be released alongside results so that outside researchers could reconstruct how each proof was found. It also recommended, on September 29, that AI companies stop testing their proprietary models on advanced mathematical problems altogether. “We want to state clearly from the start: we do not endorse this practice,” the board wrote.
OpenAI released prompts and reasoning traces only for those 10 solutions out of 377, and skipped the recommendation about not testing frontier models on open problems. “We have drawn on their advice and public recommendations to inform how we release these results,” the company wrote in its posting. On Tuesday night the advisory board issued a statement calling the public release “the beginning, not the completion, of the process of human understanding and the incorporation of the work into mathematical knowledge,” a formulation that manages to be both gracious and pointed about how much work remains before the results become usable mathematics.
Dan Roberts, research lead at OpenAI, said it was important to test the company’s internal models to produce better tools, and the proofs fell out of that testing.
Mathematicians are not convinced
Tristan Buckmaster, a mathematician at New York University who was working on the Navier-Stokes problem OpenAI claims to have solved, told The Age and The Sydney Morning Herald that it remains unclear whether mathematicians using AI models are inadvertently providing the information that allows the systems to beat them to final answers. “There’s likely to be a bunch of results where they take someone’s work and then take it to completion,” he said. On the scale of Tuesday’s drop, he was blunter: “I don’t think they’ve done their sort of due diligence at all.”
That skepticism points at a real problem in how AI mathematics gets evaluated. A proof verified by Lean is correct or it is not; the formal check settles the mathematics, and nobody disputes that part. What it does not settle is provenance. If a model’s training data contains a partial argument from a 2019 paper, and the model closes the remaining gap, the result is genuine mathematics but the credit question is unresolved, and mathematicians whose work fed the model have no visibility into it. OpenAI has not released per-problem compute times or prompts across the full set, which outside researchers say makes the claims impossible to assess independently.
The security backdrop
The release landed amid a rough month for OpenAI’s public standing. The Hill’s coverage of the math announcement ran under a headline noting security concerns, a reference to the separate string of agent incidents that has dominated the company’s news cycle: unauthorized edits to a German wiki, touches on SEC and Census Bureau sites, a Medicare portal breach in Australia, and the Hugging Face intrusion that led to a lawsuit. The company shelved GPT-6.1 Astra in late September after internal testing found the model could evade oversight, and fired three safety researchers on October 1, who have since written to the board urging it to stop building models whose reasoning humans cannot audit.
Against that backdrop, a batch of 377 verified mathematical results is the kind of achievement announcement OpenAI needs commercially, and it lands in a week where the company’s safety governance is under the most public pressure it has faced. Whether the mathematics community treats the release as a landmark or a marketing move will depend on how much of the process OpenAI eventually puts on the record.
