OpenAI math proofs fall short of academic standards
OpenAI's release of hundreds of advanced mathematics proofs has drawn criticism from elite researchers who warn that the proprietary models fail to meet rigorous academic standards.

OpenAI recently published hundreds of proposed solutions to complex mathematical problems, but the release has faced swift pushback from the academic community. Although the artificial intelligence lab consulted the Advisory Group on Mathematics and Artificial Intelligence (AGMAI)—a nine-member panel hosted by Princeton University’s Institute for Advanced Studies—the company failed to comply with several of the group's core recommendations. Most notably, AGMAI explicitly requested that labs stop testing advanced math problems on proprietary models, a guideline OpenAI ignored by using its closed systems.
Out of the 719 manuscripts OpenAI published, only 10 included the model's internal chain of thought. Furthermore, 42 percent of the proofs did not undergo formalization, a process the advisory group recommends for any solutions that humans do not yet fully comprehend. OpenAI also neglected to provide machine-readable metadata to link its natural language explanations with formal code artifacts, making it difficult for human researchers to verify the work.
This lack of transparency is particularly problematic given the translation errors discovered in the models' outputs. A recent paper by researchers at the University of Cambridge and King's College London identified at least two distinct discrepancies between the natural language proof and the Lean programming code OpenAI generated for a problem based on the Navier-Stokes equations. Because AI models can mistranslate their own logic when converting natural language into code, experts warn these outputs cannot be trusted without traditional peer review.
For AI practitioners and researchers, these shortcomings demonstrate that raw model outputs are not yet a substitute for human mathematical rigor. Prominent mathematician Terence Tao noted online that autonomous AI prompters often lack interest in the broader field and cannot interact with other researchers to explain their results. Until AI developers prioritize human comprehension and fund collaborative verification efforts, practitioners must treat automated proofs as unverified drafts rather than established scientific breakthroughs.
This is our own summary of reporting by TechCrunch AI



