Proof or Bluff? Evaluating LLMs on 2025 USA Math Olympiad
2025/03/27 by Ivo Petrov, Petrov, Ivo, Jasper Dekoninck +13 · 35 voices · 31 citations
#cs.CL
paper · pdf · doi:10.48550/arxiv.2503.21934
Abstract
Recent math benchmarks for large language models (LLMs) such as MathArena indicate that state-of-the-art reasoning models achieve impressive performance on mathematical competitions like AIME, with the leading model, Gemini-2.5-Pro, achieving scores comparable to top human competitors. However, these benchmarks evaluate models solely based on final numerical answers, neglecting rigorous reasoning and proof generation which are essential for real-world mathematical tasks. To address this, we introduce a comprehensive evaluation of full-solution reasoning for challenging mathematical problems. Using expert human annotators, we evaluated several state-of-the-art reasoning models on the six problems from the 2025 USAMO within hours of their release. Our results reveal that all tested models struggled significantly: only Gemini-2.5-Pro achieves a non-trivial score of 25%, while all other models achieve less than 5%. Through detailed analysis of reasoning traces, we identify the most common failure modes and find several unwanted artifacts arising from the optimization strategies employed during model training. Overall, our results suggest that current LLMs are inadequate for rigorous mathematical reasoning tasks, highlighting the need for substantial improvements in reasoning and proof generation capabilities.
Citations
Cited by
Discussions
- Reminder that LLMs are a dismal expensive prototype with a handful of potential applications, not a reliable general purpose technology. Researchers tried a couple on the 2025 USA Mathematical Olympia [bsky, 251 points, 6 comments]
- arxiv.org/abs/2503.21934 tl;dr: people think LLMs are getting good at maths because they're being tested on hard questions for which the answer is a number. But when you ask them for proofs (which is [bsky, 52 points, 2 comments]
- Tests on USAMO immediately after problems were posted yield surprisingly bad model performance. Suggests there's much more training on test than expected. arxiv.org/abs/2503.219... [bsky, 31 points, 7 comments]
- yr not gonna believe this but if you give super SoTA LLMs math questions that they can't have memorized (just released math olympiad Q's) they fall entirely flat on their faces: arxiv.org/abs/2503.219 [bsky, 18 points, 3 comments]
- Proof or Bluff? Evaluating LLMs on 2025 USA Math Olympiad [hn, 13 points, 2 comments]
- AI bros want you to believe their LLMs are math prodigies. When given math problems that weren’t already online (taken from the 2025 USA Math Olympiad), they scored an average of less than 5% [bsky, 10 points, 0 comments]
- Proof or Bluff? Evaluating LLMs on 2025 USA Math Olympiad arxiv.org/abs/2503.21934 cc @edzitron.com @emilymbender.bsky.social [bsky, 8 points, 0 comments]
- Here’s a paper that when given novel math problems llms performed poorly. [bsky, 8 points, 2 comments]
- Proof or Bluff? Evaluating LLMs on 2025 USA Math Olympiad [hn, 6 points, 1 comments]
- Would also add: Reasoners excel in STEM because of large amount of training data with very similar content. Reasoners still continue to fail in STEM when faced with new content (see recent results on [bsky, 6 points, 0 comments]
- Proof or Bluff? Evaluating LLMs on 2025 USA Math Olympiad [hn, 5 points, 0 comments]
- New study shows why simulated reasoning AI models don’t yet live up to their billing [lemmy, 3 points, 0 comments]
- Proof or Bluff? Evaluating LLMs on 2025 USA Math Olympiad [hn, 3 points, 1 comments]
- Proof or Bluff? Evaluating LLMs on 2025 USA Math Olympiad [hn, 3 points, 0 comments]
- Proof or Bluff? Evaluating LLMs on 2025 USA Math Olympiad [hn, 3 points, 0 comments]
- reminds me a lot of these results arxiv.org/abs/2503.21934 [bsky, 2 points, 0 comments]
- arxiv.org/abs/2503.21934 Las matemáticas siguen siendo el muro de hielo de la IAGen. Todos los modelos se comportaron de forma mediocre respondiendo a las preguntas de las Olimpiadas Matemáticas de es [bsky, 2 points, 3 comments]
- Proof or Bluff? Evaluating LLMs on 2025 USA Math Olympiad [hn, 2 points, 0 comments]
- Despite popular belief it is not true that current LLMs can solve math. olympiad problems. (I have a set of Czech middle school problems Klokánek, most LLMs are slightly above random chance, best thin [bsky, 2 points, 1 comments]
- arxiv.org/abs/2503.219... www.reddit.com/r/LocalLLaMA... Člověk by skoro až řekl, že si prostě LLM pamatují stará řešení, ale na letošní zadání zírají jak tele na nová vrata? [bsky, 2 points, 0 comments]
- Proof or Bluff? Evaluating LLMs on 2025 USA Math Olympiad [hn, 2 points, 0 comments]
- Proof or Bluff? Evaluating LLMs on 2025 USA Math Olympiad [hn, 2 points, 0 comments]
- How did LLMs perform hours after the 2025 US Math Olympiad's problems were released? Well, they sucked, scoring 5% on average. Who would have thought? arxiv.org/abs/2503.219... cc @edzitron.com [bsky, 1 points, 0 comments]
- Gli LLM più evoluti messi davanti a un problema che non è nel loro addestramento, in questo caso le domande delle olimpiadi di matematica 2025 appena uscite, fanno una figuraccia, come prevedibile. Te [bsky, 1 points, 0 comments]
- A cold shower by INSAIT folks for the much hyped reasoning in the top current LLMs. When subjected to scrutiny the thinking in these models reveals that the way they arrive at answers that require act [bsky, 1 points, 0 comments]
- I finally read the details in this paper assessing mathematical reasoning in LLMs. They test competence rather than performance (show your work rather than multiple choice) and...5% success rate for a [bsky, 1 points, 1 comments]
- “Current LLMs struggle significantly on USA Math Olympiad problems, with the best-performing model achieving an average score of less than 25% […] all other models achieve less than 5%.” They don’t th [bsky, 1 points, 0 comments]
- So AIs are dilettantes, they know everything but understand nothing Elevating AI performance on reasoning tasks Commentary here: arstechnica.com/ai/2025/04/n... arxiv.org/abs/2503.21934 [bsky, 1 points, 0 comments]
- Proof or Bluff? Evaluating LLMs on 2025 USA Math Olympiad [hn, 1 points, 0 comments]
- www.arxiv.org/abs/2503.21934 [bsky, 0 points, 0 comments]
- Link to paper: arxiv.org/pdf/2503.219... [bsky, 0 points, 0 comments]
- "we evaluated several state-of-the-art reasoning models on the six problems from the 2025 USA Mathematical Olympiad within hours of their release. Our results reveal that all tested models struggled s [bsky, 0 points, 0 comments]
- I prefer “intuitively obvious,” but “trivial” is still some good hand-waving. arxiv.org/abs/2503.21934 [bsky, 0 points, 0 comments]
- Proof or Bluff paper from ETH Zurich, INSAIT shows evaluation improvements in LLMs may be a mirage, frontier models struggle with mathematical reasoning, particularly in proof generation for complex p [bsky, 0 points, 0 comments]
- Who could have predicted this? 🙄 state-of-the-art LLMs score 5% on the 2025 mathematical olympiad despite having been trained extensively on past editions : https:// arxiv.org/abs/2503.21934 # ai # A [mastodon, 0 points, 0 comments]
Related