The Leaderboard Illusion
2025/04/29 by Shivalika Singh, Nan Yu, Singh, Shivalika +26 · 32 voices · 20 citations
Medicine · Computer Science · Social Sciences · #Artificial Intelligence in Healthcare and Education #AI in Service Interactions #Ethics and Social Impacts of AI
paper · pdf · doi:10.48550/arxiv.2504.20879
Abstract
Measuring progress is fundamental to the advancement of any scientific field. As benchmarks play an increasingly central role, they also grow more susceptible to distortion. Chatbot Arena has emerged as the go-to leaderboard for ranking the most capable AI systems. Yet, in this work we identify systematic issues that have resulted in a distorted playing field. We find that undisclosed private testing practices benefit a handful of providers who are able to test multiple variants before public release and retract scores if desired. We establish that the ability of these providers to choose the best score leads to biased Arena scores due to selective disclosure of performance results. At an extreme, we identify 27 private LLM variants tested by Meta in the lead-up to the Llama-4 release. We also establish that proprietary closed models are sampled at higher rates (number of battles) and have fewer models removed from the arena than open-weight and open-source alternatives. Both these policies lead to large data access asymmetries over time. Providers like Google and OpenAI have received an estimated 19.2% and 20.4% of all data on the arena, respectively. In contrast, a combined 83 open-weight models have only received an estimated 29.7% of the total data. We show that access to Chatbot Arena data yields substantial benefits; even limited additional data can result in relative performance gains of up to 112% on the arena distribution, based on our conservative estimates. Together, these dynamics result in overfitting to Arena-specific dynamics rather than general model quality. The Arena builds on the substantial efforts of both the organizers and an open community that maintains this valuable evaluation platform. We offer actionable recommendations to reform the Chatbot Arena's evaluation framework and promote fairer, more transparent benchmarking for the field
Cited by
Discussions
- The Leaderboard Illusion [hn, 184 points, 51 comments]
- So, your favorite/fancy/rich AI provider most likely cheats most of the time to score high in LLMs leaderboards. Shocking, but totally expected, isn't it? arxiv.org/abs/2504.20879 #AI #LLM [bsky, 51 points, 4 comments]
- We tried very hard to get this right, and have spent the last 5 months working carefully to ensure rigor. If you made it this far, take a look at the full 68 pages: arxiv.org/abs/2504.20879 Any feedba [bsky, 12 points, 1 comments]
- "When a metric becomes a target it ceases to be a useful metric." arxiv.org/abs/2504.20879 [bsky, 9 points, 0 comments]
- 📖For today's Reading Group Joe Baumann presented "The Leaderboard Illusion" by Shivalika Singh et al. Paper: arxiv.org/pdf/2504.20879 #NLProc [bsky, 7 points, 0 comments]
- Some interesting detective work alleging that proprietary LLM developers are gaming the Chatbot Arena leaderboards, with collusion from Chatbot Arena's operators. [bsky, 7 points, 2 comments]
- This would be the Scandal of the Century if it happened in the chess world! Imagine a player secretly playing multiple matches against an opponent but reporting only the score from the best match to m [bsky, 7 points, 1 comments]
- First up is @mziizm.bsky.social, discussing "mirrors" and evaluation. Most evaluation measures correlation with Arena rankings, but "when a measure becomes a target, it ceases to be a good measure." S [bsky, 5 points, 1 comments]
- wenn du nachsehen würdest, könntest du durchaus zu dem Schluss kommen, ja. Sind journalist founded, ehemals WIRED. Aber das zu LMArena haben sie sich ja nicht ausgedacht, ist aus nem Forschungspapier [bsky, 2 points, 0 comments]
- The Leaderboard Illusion "these dynamics result in overfitting to Arena-specific dynamics rather than general model quality." arxiv.org/abs/2504.2... #GenerativeAI #ArtificialIntelligence [bsky, 2 points, 0 comments]
- >> Read the paper in full: arxiv.org/pdf/2504.20879 @sarahooker.bsky.social @cohereforai.bsky.social @mziizm.bsky.social @sayash.bsky.social @shaynelongpre.bsky.social @beyzaermis.bsky.social @sanmiko [bsky, 2 points, 0 comments]
- Researchers report Chatbot Arena scores are flawed. An undisclosed policy allows preferred providers disproportionate private testing access, making the leaderboard an unreliable measure of model capa [bsky, 1 points, 0 comments]
- Stanford, Princeton, MIT et al have released a research paper that challenges the illusion of the AI leaderboards, covering some of the issues as well as some concrete recommendations. Paper: arxiv.or [bsky, 1 points, 0 comments]
- Big AI companies are getting help behind the scenes from the people that run the tests. Fuck that. [bsky, 0 points, 0 comments]
- Sad to see, but true: LLM companies are systematically gaming leaderboards, to make LLMs seem smarter than they are. This is why the Apple paper from 2 weeks ago had so much impact: it is part of a st [bsky, 0 points, 0 comments]
- The Leaderboard Illusion https://arxiv.org/abs/2504.20879 [bsky, 0 points, 0 comments]
- 8/ Leaderboards are powerful, but power demands scrutiny. Chatbot Arena is a vital community resource. Let’s make it fair, transparent, and scientifically sound. Read the full paper here: arxiv.org/ab [bsky, 0 points, 1 comments]
- Paper reveals bias in Chatbot Arena: closed AI providers get unfair advantages—more data, more control, & selective score reporting. Open models get sidelined. The current leaderboard fosters overfitt [bsky, 0 points, 0 comments]
- Super interesting paper on how current ML evaluation rankings are gamed - benchmarks and evaluation are important topics: arxiv.org/abs/2504.20879 [bsky, 0 points, 1 comments]
- Researchers Say the Most Popular Tool for Grading AIs Unfairly Favors Meta, Google, OpenAI [lemmy, 0 points, 0 comments]
- The Leaderboard Illusion https://arxiv.org/abs/2504.20879 (https://news.ycombinator.com/item?id=43842380) [bsky, 0 points, 0 comments]
- The Leaderboard Illusion https://arxiv.org/abs/2504.20879 (https://news.ycombinator.com/item?id=43842380) [bsky, 0 points, 0 comments]
- The leaderboard illusion: arxiv.org/abs/2504.20879 Look at some of the shortcomings of the Chatbot arena which has emerged as the go-to leaderboard for ranking the most capable #AI . Recently Meta abu [bsky, 0 points, 0 comments]
- Looks like LLM vendors are aggressively gaming leaderboards, not surprising considering the large amounts of money at stake... But still seems to be little interest in moving to more real-world evalua [bsky, 0 points, 0 comments]
- Classer les IA générative n'est il pas une illusion pour estimer leur valeur. ? arxiv.org/abs/2504.20879 [bsky, 0 points, 0 comments]
- The Leaderboard Illusion [bsky, 0 points, 0 comments]
- The Leaderboard Illusion #HackerNews https://arxiv.org/abs/2504.20879 [bsky, 0 points, 0 comments]
- The Leaderboard Illusion https://arxiv.org/abs/2504.20879 https://news.ycombinator.com/item?id=43842380 [bsky, 0 points, 0 comments]
- The Leaderboard Illusion https://arxiv.org/abs/2504.20879 [bsky, 0 points, 0 comments]
- Post: The Leaderboard Illusion View Article | Join the HN Conversation Thread with summary 🧵👇 #hacker-news [bsky, 0 points, 1 comments]
- The Leaderboard Illusion #llms #leaderboard #benchmark #arena #evaluation [bsky, 0 points, 0 comments]
- I have used and contributed to lmsys loads of times; such a disappointment that the "leading" models on lmarena.ai are likely being gamed. Does not feel scientific, feels dishonest. SOURCE: arxiv.org/ [bsky, 0 points, 0 comments]
Related