The Wire
TechnologyArtificial IntelligenceScience & Health

OpenAI releases 722 math manuscripts, but verification remains uneven

OpenAI releases 722 math manuscripts, but verification remains uneven
Photo: theverge.com

OpenAI released 722 mathematical manuscripts grouped into 372 result families.

Why it matters: The collection could expand AI-assisted mathematical research if its claims are independently validated. But the cited release materials do not establish that every result is correct, novel, publication-ready or reproducible.

  • The repository snapshot reviewed lists 722 manuscripts organized into 372 result families across 17 broad mathematical areas.
  • OpenAI's repository documentation says most results came from roughly 4,000 research problems and used about three hours of ChatGPT Pro-equivalent compute per retained result.
  • A third-party inventory reports 162 manuscripts explicitly listing a Lean formalization of the main result, or 162 divided by 722, equal to 22.4%.
  • The cited release materials do not report the model weights, full prompts, sampling settings, training data or complete compute logs as available.

OpenAI says it is releasing a broad collection of mathematical results produced by an internal frontier model. The repository documentation lists 722 manuscripts organized into 372 result families across 17 broad areas. These are snapshot figures; the supplied materials do not state an access date, and the repository may change.

The README presents the 372 families as the repository's grouping, but the cited material does not define the grouping criteria. It therefore does not establish that the families represent 372 independent breakthroughs. Companion arguments, consequences or alternative proofs may appear within a family, but that interpretation is not defined in the cited release documentation.

OpenAI's repository documentation says most results came from an evaluation involving approximately 4,000 research problems and that each retained result used roughly three hours of ChatGPT Pro-equivalent thinking compute on average. The collection includes claims involving long-standing mathematics problems, according to The Verge.

Verification is partial. OpenAI says many, but not all, manuscripts have Lean formalizations and warns that “some of the unformalized results could have issues,” according to the repository documentation.

A third-party inventory counted 162 manuscripts whose entries explicitly listed a Lean formalization of the main result. Its access date is not stated in the supplied materials. The calculation here is manuscript-level: 162 divided by the repository's 722-manuscript count equals 22.4%, not a share of the 372 result families.

Lean can check a specified formal statement and proof artifact within a formal system. It does not by itself establish that an informal claim is important, novel, correctly stated, properly attributed or independently validated, as another repository analysis notes.

The cited release does not report a complete independent audit of all 722 manuscripts or 372 families. Nor do the release materials cited here report journal submissions, acceptances, rejections or external peer review. They also do not report the model weights, full prompts, sampling settings, training data or complete compute logs as available, so the materials cited here do not establish full reproducibility.

By the numbers

  • 722 - manuscripts listed in the repository snapshot reviewed
  • 372 - result families listed by the repository
  • 22.4% - 162 manuscripts with an explicitly listed Lean formalization, divided by 722

Yes, but: The figures are time-sensitive repository counts, and the supplied materials do not provide access dates for the repository or third-party inventory. Lean formalization also checks only the specified formal statement and proof artifact, not every aspect of an informal research claim.

Based on reporting from

  • The Verge

See how this story touches your network - open The Wire in Jane.

Open in Jane