OpenAI posts a new batch of math results on GitHub

By

Published

Reporting from The Verge, OpenAI, WIRED

OpenAI released a new batch of AI math results on GitHub with Lean proofs, as mathematicians say their advice was ignored.

What it means for founders

  • Verification is the real cost. A pile of 722 manuscripts is cheap to publish and expensive to check. Machine-checkable proofs like Lean are what make claims credible at scale; if you sell AI research or analysis, ship evidence a customer can verify, not a press release.
  • Compute disclosure gives you a yardstick. About three hours of ChatGPT Pro reasoning per result is a rough unit for pricing deep-reasoning work. Watch whether the released model matches it.
  • Trust is now a platform risk. The labs your product depends on are in open dispute with domain experts over disclosure and credit. Expect technical buyers to ask where AI outputs came from and who checked them.
  • What to watch: AGMAI's verdict on this batch, any move of the papers to Hexagon or Palomar, the dates of OpenAI's promised workshops, and a release date for the model.

The story

OpenAI on October 6 released a new batch of math results produced by an unreleased internal frontier model, posting the papers to GitHub and describing the approach in a blog post. The Verge counted 722 manuscripts grouped into 372 result families. AGMAI, an independent advisory group of mathematicians linked to the Institute for Advanced Study, says the batch resolves hundreds of open questions. Those are claims: the wider field has not yet worked through the papers.

This release is separate from OpenAI's earlier Navier-Stokes claim. OpenAI says it drew on advice from that independent math panel in deciding how to publish.

What is in the OpenAI math release

  • Papers on GitHub, with protocols for revisions and citations. OpenAI says it is still looking at community-hosted alternatives that meet the advisory group's guidelines.
  • Lean formalizations for many of the proofs, so a computer can check them. OpenAI says it will add more as it gets them.
  • Process data: 10 write-ups describing how the model reasoned, compute estimates expressed as ChatGPT Pro usage, and counts of attempted problems. OpenAI says the average result took compute equal to about three hours of thinking time on ChatGPT Pro.
  • Funding pledges for workshops and conferences on AI-produced results, plus a stated plan to release the model responsibly, with no date.

AGMAI's late September recommendations asked labs to use normal academic venues when they can, disclose the model, prompts and compute, and stop treating math results as marketing.

Mathematicians say their advice was ignored

Before the release, WIRED reported that around 40 mathematicians met OpenAI in August and asked it to publish explanatory papers rather than announce results by blog or tweet. Northwestern mathematician Bryna Kra told the magazine: "Apparently, that input was ignored." Kra also said OpenAI has been pointed to community tools such as Hexagon, which mainly hosts AI-generated work, and Palomar, which records machine-verified proofs, without changing its behavior.

Attendees told WIRED the company had assured them it would not release everything at once; OpenAI spokesperson Lindsay McCallum said OpenAI has no knowledge of such a promise. Nestor Guillen, a visiting professor at NYU, described a perception of mobster-like conduct by AI companies. McCallum said OpenAI disagrees with that characterization and is working with mathematicians on the field's future.

What we don't know yet

  • How many of the results survive expert review, and which proofs lack a Lean formalization.
  • Whether the papers move to a community-hosted venue, as AGMAI recommends.
  • Whether AGMAI will say publicly if this release met its guidelines.
  • When, or whether, the model behind the results becomes available.

Sources

Primary sources

Enki Daily

Get stories like this every weekday morning.

The day's AI stories for founders, each with what it means for your company. Free.

More in Research

How Enki covers newsCorrectionsReport an error

Search Enki

Search AI tools, categories and news