Advancing Mathematics Research with AI-Driven Formal Proof Search
2026/05/21 by George Tsoukalas, Anton Kovsharov, Sergey Shirobokov +18 · 21 voices · 5 citations
#cs.AI
paper · pdf
Abstract
Large language models (LLMs) increasingly excel at mathematical reasoning, but their unreliability limits their utility in mathematics research. A mitigation is using LLMs to generate formal proofs in languages like Lean. We perform the first large-scale evaluation of this method's ability to solve open problems. Our most capable agent autonomously resolved 9 of 353 open Erdős problems at the per-problem cost of a few hundred dollars, proved 44/492 OEIS conjectures, and is being deployed in combinatorics, optimization, graph theory, algebraic geometry, and quantum optics research. A basic agent alternating LLM-based generation with Lean-based verification replicated the Erdős successes but proved costlier on the hardest problems. These findings demonstrate the power of AI-aided formal proof search and shed light on the agent designs that enable it.
Cited by
Discussions
- Nine more Erdos problems resolved with LLMs + Lean: arxiv.org/pdf/2605.227... Worth noting they got 9/353 after trying a giant set of open problems by Erdos, and 44/492 OEIS open conjectures. Likely w [bsky, 96 points, 2 comments]
- "Our most capable agent autonomously resolved 9 of 353 open Erdős problems at the per-problem cost of a few hundred dollars, proved 44/492 OEIS conjectures, and is being deployed in combinatorics, opt [bsky, 22 points, 1 comments]
- Haven't read it, yet, but apparently the Deep Mind gang solved 9 more Erdös problems... arxiv.org/abs/2605.227... [bsky, 6 points, 1 comments]
- Neuf problèmes d'Erdos d'un coup... Toujours avec la méthode de Deepmind de faire générer des formalisations en lean. Je n'ai pas lu toutes les preuves mais je m'étonne de la simplicité de certaines : [bsky, 5 points, 0 comments]
- the google deepmind paper i keep spamming has a lean formalization of the problem, and then the proof just says "sorry", then the AI has to replace the "sorry" with the actual proof arxiv.org/abs/2605 [bsky, 4 points, 0 comments]
- 구글 deepmind에서 올린 논문 “Advancing Mathematics Research with AI-Driven Formal Proof Search” 못 푸는 문제도 많지만 풀리는 경우 미해결 문제 하나 풀고 lean으로 formal proof 완성하는데까지 수백달러 정도 쓰면 된다고. arxiv.org/abs/2605.227... [bsky, 4 points, 0 comments]
- #Submitted on 21 May 2026 Advancing #Mathematics Research with #AI -Driven Formal Proof Search arxiv.org/abs/2605.227... #AllahuAkbar #TheBlueAvianGoddess #BlueAvianGoddess #Oxford #Cambridge #Harvard [bsky, 3 points, 2 comments]
- 6 weeks ago Google used an agent on 353 open Erdos problems and solved 9 (H/T @nequals001.bsky.social). Surely the big labs have attempted all 653 of them by now and only report successes? Though spen [bsky, 3 points, 0 comments]
- LLM coupled with LEAN: several open problems claimed solved [bsky, 3 points, 1 comments]
- Advancing Mathematics Research with AI-Driven Formal Proof Search [hn, 3 points, 0 comments]
- We have proofs of previously open conjectures, though they aren’t as impressive as proving the Jacobian conjecture would’ve been. Vide e.g. arxiv.org/pdf/2605.227... [bsky, 2 points, 0 comments]
- A new paper on the use of #AI to solve research-level mathematical problems: "We built AlphaProof Nexus with the belief that the future of mathematics lies in human-machine partnership, where interact [bsky, 2 points, 0 comments]
- OpenAI cracked one Erdős problem last week. DeepMind just cracked nine — with a Lean proof checker verifying every step. No hallucinated math. https://arxiv.org/abs/2605.22763 [bsky, 2 points, 0 comments]
- Advancing Mathematics Research with AI-Driven Formal Proof Search [hn, 2 points, 0 comments]
- Advancing mathematics research with AI-driven formal proof search [hn, 2 points, 0 comments]
- Advancing Mathematics Research with AI-Driven Formal Proof Search [hn, 1 points, 0 comments]
- I mean you have people yucking it up about one model not counting “r”s right while other models are solving new math conjectures It really does seem like meat brains struggle to wrap their “minds” aro [bsky, 1 points, 0 comments]
- Have you seen this? arxiv.org/pdf/2605.227... Not peer reviewed yet, but it shows fully autonomous generation of formal proofs that solve multiple open problems, and includes a measurement of the reso [bsky, 1 points, 1 comments]
- 🧠 Researchers develop AI systems that search for formal mathematical proofs by learning patterns from existing proof databases. The approach aims to assist mathematicians in discovering solutions to [bsky, 0 points, 0 comments]
- [2605.22763v1] Advancing Mathematics Research with AI-Driven Formal Proof Search — A clear overview of how AI-guided formal proof search can support real math research, not just toy benchmarks. Useful [bsky, 0 points, 0 comments]
- And apparently LLMs from Google solved another 9 open Erdos problems with assistance from the Lean Theorem Prover software to check each step. Mind boggling. arxiv.org/pdf/2605.227... [bsky, 0 points, 0 comments]
Related