Skip to main content

ITmatterss

Even AI struggles with maths: OpenAI’s latest proofs face scrutiny

Vertical Share Bar
Even AI struggles with maths: OpenAI’s mathematical proofs face scrutiny

News in Short

  • OpenAI released 719 mathematical manuscripts containing solutions to difficult research problems.
  • Researchers have questioned whether the proofs meet standards for human understanding and verification.
  • Only 42% of the released proofs had been formally verified, according to the report.
  • A separate paper identified discrepancies between a natural-language proof and its Lean code.
  • Mathematicians say AI-generated solutions still need peer review and human scrutiny before researchers can trust them.

Artificial intelligence can write code, analyse data and tackle complex reasoning tasks. However, OpenAI’s latest mathematical work highlights a persistent challenge: producing solutions that researchers can independently understand and verify.

The company recently released 719 mathematical manuscripts containing AI-generated proofs. The effort aimed to demonstrate how its models could tackle some of the hardest problems in mathematics.

However, the release has raised questions about whether these results meet the standards expected by mathematicians. Researchers have pointed to gaps in formal verification, limited transparency and discrepancies between written explanations and computer-checkable proofs.

The debate highlights an important distinction. An AI model may produce a plausible mathematical solution, but that does not automatically make the result trustworthy or useful to the wider research community.

OpenAI’s mathematical proofs face scrutiny

OpenAI consulted an advisory group of prominent mathematicians before releasing its latest work. The group, known as the Advisory Group on Mathematics and Artificial Intelligence (AGMAI), published guidelines for AI companies working on difficult mathematical problems in September.

However, the group’s recommendations raise questions about the latest release. One recommendation was to stop testing advanced mathematical problems on proprietary models. OpenAI’s release explicitly states that it evaluates its proprietary models using open research problems.

The release also offered different levels of transparency across its manuscripts. According to the report, only 10 of the 719 manuscripts included the models’ chain of thought.

AGMAI’s guidelines also emphasised formal verification when humans cannot readily understand a proof. Yet only 42% of OpenAI’s released proofs had undergone formalisation, according to the report.

These figures do not establish that the remaining proofs are incorrect. Instead, they highlight how much work may remain before the mathematical community can independently assess the results.

AI-generated proofs can lose meaning during translation

One of the biggest concerns involves how AI models convert mathematical reasoning into a verifiable format.

Models often generate a natural-language explanation first. They then translate that explanation into Lean, a programming language used to check mathematical proofs formally.

In principle, Lean can help verify whether a proof follows the required logical rules. However, the translation itself can introduce problems.

A paper from researchers at the University of Cambridge and King’s College London identified at least two discrepancies between a natural-language proof and its Lean code for a problem derived from the Navier-Stokes equations.

These equations describe complex fluid behaviour. The researchers’ findings do not necessarily invalidate either solution. However, they raise questions about whether the formal code accurately represents the reasoning explained in the original proof.

This matters because a formally checked proof is only useful if it corresponds to the mathematical claim being made. If the explanation and code differ, researchers need to investigate the discrepancy rather than assume that verification settles the matter.

Why human mathematicians still matter

Mathematical research involves more than reaching an answer. Researchers must explain their reasoning, defend their methods and help others understand why a result matters.

Human mathematicians typically publish papers, present their findings and respond to questions from peers. That process helps uncover errors and can reveal techniques that other researchers can apply to different problems.

AI-generated results introduce a different challenge. A model can produce a proposed solution without a human researcher fully understanding every step.

Mathematician Terence Tao has criticised approaches in which people prompt AI to solve difficult problems but cannot adequately explain the resulting work. Harvard mathematics professor Melanie Wood has similarly emphasised that human understanding often begins after an AI produces its result.

Consequently, the burden does not end when a model generates a proof. Researchers must still examine it, verify its logic and determine whether it advances mathematical knowledge.

AI’s maths challenge is about trust, not just accuracy

OpenAI’s latest release illustrates why mathematical reasoning remains a demanding test for AI systems. Producing a convincing answer is one challenge. Demonstrating that the answer is correct, reproducible and understandable is another.

Formal verification can help, but it does not eliminate every problem. Researchers must also establish that the formal representation accurately reflects the original argument.

AGMAI has called for machine-readable metadata linking natural-language proofs with their formal counterparts. The aim is to make it easier for researchers to compare both versions and identify inconsistencies.

The group has also suggested that AI companies help fund the human work required to evaluate and understand their results.

For OpenAI, the next challenge is therefore not simply to produce more mathematical solutions. It is to ensure that independent researchers can validate them and build on them.

AI may increasingly help mathematicians explore difficult problems. However, until its solutions withstand rigorous scrutiny, a model’s claim to have solved a problem should be treated as a starting point rather than the final word.

41

Leave a Reply

Your email address will not be published. Required fields are marked *

logo

Get the latest news instantly

You can change your preferences anytime.