Towards Contamination Resistant Benchmarks
2025/05/13 by Musawi, Rahmatullah, Lu, Sheng
#Computation and Language (cs.CL) #FOS: Computer and information sciences
paper · doi:10.48550/arxiv.2505.08389
Abstract
The rapid development of large language models (LLMs) has transformed the landscape of natural language processing. Evaluating LLMs properly is crucial for understanding their potential and addressing concerns such as safety. However, LLM evaluation is confronted by various factors, among which contamination stands out as a key issue that undermines the reliability of evaluations. In this work, we introduce the concept of contamination resistance to address this challenge. We propose a benchmark based on Caesar ciphers (e.g., "ab" to "bc" when the shift is 1), which, despite its simplicity, is an excellent example of a contamination resistant benchmark. We test this benchmark on widely used LLMs under various settings, and we find that these models struggle with this benchmark when contamination is controlled. Our findings reveal issues in current LLMs and raise important questions regarding their true capabilities. Our work contributes to the development of contamination resistant benchmarks, enabling more rigorous LLM evaluation and offering insights into the true capabilities and limitations of LLMs.
Citations
- Leveraging Online Olympiad-Level Math Problems for LLMs Training and Contamination-Resistant Evaluation
- Qwen2.5 Technical Report
- LiveXiv -- A Multi-Modal Live Benchmark Based on Arxiv Papers Content
- When a language model is optimized for reasoning, does it still show embers of autoregression? An analysis of OpenAI o1
- Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters
- A Careful Examination of Large Language Model Performance on Grade School Arithmetic
- LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code
- TRUCE: Private Benchmarking to Prevent Contamination and Improve Comparative Evaluation of LLMs
- Generalization or Memorization: Data Contamination and Trustworthy Evaluation for Large Language Models
- Leak, Cheat, Repeat: Data Contamination and Evaluation Malpractices in Closed-Source LLMs
- Task Contamination: Language Models May Not Be Few-Shot Anymore
- How are Prompts Different in Terms of Sensitivity?
- In-context Learning Generalizes, But Not Always Robustly: The Case of Syntax
- NLP Evaluation in trouble: On the Need to Measure LLM Data Contamination for each Benchmark
- Are Emergent Abilities in Large Language Models just In-Context Learning?
- Instructed to Bias: Instruction-Tuned Language Models Exhibit Emergent Cognitive Bias
- Trained Transformers Learn Linear Models In-Context
- Self-contradictory Hallucinations of Large Language Models: Evaluation, Detection and Mitigation
- Language Models Implement Simple Word2Vec-style Vector Arithmetic
- Sources of Hallucination by Large Language Models on Inference Tasks
- Adaptive Chameleon or Stubborn Sloth: Revealing the Behavior of Large Language Models in Knowledge Conflicts
- How Language Model Hallucinations Can Snowball
- Sparks of Artificial General Intelligence: Early experiments with GPT-4
- A Multitask, Multilingual, Multimodal Evaluation of ChatGPT on Reasoning, Hallucination, and Interactivity
- Emergent Analogical Reasoning in Large Language Models
- Can In-context Learners Learn a Reasoning Concept from Demonstrations?
- What learning algorithm is in-context learning? Investigations with linear models
- Scaling Instruction-Finetuned Language Models
- Language Models are Multilingual Chain-of-Thought Reasoners
- Language Models Are Greedy Reasoners: A Systematic Formal Analysis of Chain-of-Thought
- WinoDict: Probing language models for in-context word acquisition
- What Can Transformers Learn In-Context? A Case Study of Simple Function Classes
- Emergent Abilities of Large Language Models
- Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models
- Large Language Models are Zero-Shot Reasoners
- PaLM: Scaling Language Modeling with Pathways
- BNAI, NO-TOKEN, and MIND-UNITY: Pillars of a Systemic Revolution in Artificial Intelligence
- Training Verifiers to Solve Math Word Problems
- Prompt Programming for Large Language Models: Beyond the Few-Shot Paradigm
- Language Models are Few-Shot Learners
- SuperGLUE: A Stickier Benchmark for General-Purpose Language Understanding Systems
- Know What You Don't Know: Unanswerable Questions for SQuAD
- LiveBench: A Challenging, Contamination-Limited LLM Benchmark
- The Surprising Effectiveness of Test-Time Training for Few-Shot Learning
- Language Ranker: A Metric for Quantifying LLM Performance Across High and Low-Resource Languages
- Are Emergent Abilities of Large Language Models a Mirage?
- Why Can GPT Learn In-Context? Language Models Implicitly Perform Gradient Descent as Meta-Optimizers
- Large Language Models Are Latent Variable Models: Explaining and Finding Good Demonstrations for In-Context Learning
Related