How Should AI Safety Benchmarks Benchmark Safety?
2026/01/31 by Cheng Yu, Severin Engelmann, Ruoxuan Cao +2 · 3 citations
Computer Science · #cs.CY
paper · pdf · doi:10.48550/arxiv.2601.23112
arxiv created 2026/08/05 · arxiv updated 2026/08/06
Abstract
AI safety benchmarks are pivotal for safety in advanced AI systems; however, they have significant technical, epistemic, and sociotechnical shortcomings. We present a review of 210 safety benchmarks that maps out common challenges in safety benchmarking, documenting failures and limitations by drawing from engineering sciences and long-established theories of risk and safety. We argue that adhering to established risk management principles, mapping the space of what can(not) be measured, developing robust probabilistic metrics, and efficiently deploying measurement theory to connect benchmarking objectives with the world can significantly improve the validity and usefulness of AI safety benchmarks. The review provides a roadmap on how to improve AI safety benchmarking, and we illustrate the effectiveness of these recommendations through quantitative and qualitative evaluation. We also provide workflow-oriented guiding questions with illustrative benchmark that help researchers and practitioners develop robust and epistemologically sound safety benchmarks. This study advances the science of benchmarking and helps practitioners deploy AI systems more responsibly.
Citations
- Open-World Deepfake Attribution via Confidence-Aware Asymmetric Learning
- The Role of Risk Modeling in Advanced AI Risk Management
- Operationalizing Pluralistic Values in Large Language Model Alignment Reveals Trade-offs in Safety, Inclusivity, and Model Behavior
- Measuring what Matters: Construct Validity in Large Language Model Benchmarks
- Risk Management for Mitigating Benchmark Failure Modes: BenchRisk
- Measuring What Matters: Connecting AI Ethics Evaluations to System Attributes, Hazards, and Harms
- QA-LIGN: Aligning LLMs through Constitutionally Decomposed QA
- Systematic Hazard Analysis for Frontier AI using STPA
- CoP: Agentic Red-teaming for Large Language Models using Composition of Principles
- B-score: Detecting biases in large language models using response history
- MDIT-Bench: Evaluating the Dual-Implicit Toxicity in Large Multimodal Models
- Measurement to Meaning: A Validity-Centered Framework for AI Evaluation
- What Is AI Safety? What Do We Want It to Be?
- Adapting Probabilistic Risk Assessment for AI
- PROMPTEVALS: A Dataset of Assertions and Guardrails for Custom Production Large Language Model Pipelines
- The Hidden Space of Safety: Understanding Preference-Tuned LLMs in Multilingual context
- Cultural Learning-Based Culture Adaptation of Language Models
- KOFFVQA: An Objectively Evaluated Free-form VQA Benchmark for Large Vision-Language Models in the Korean Language
- Proof or Bluff? Evaluating LLMs on 2025 USA Math Olympiad
- FLEX: A Benchmark for Evaluating Robustness of Fairness in Large Language Models
- The Greatest Good Benchmark: Measuring LLMs' Alignment with Utilitarian Moral Dilemmas
- When Tom Eats Kimchi: Evaluating Cultural Bias of Multimodal Large Language Models in Cultural Mixture Contexts
- Social Bias Benchmark for Generation: A Comparison of Generation and QA-Based Evaluations
- AILuminate: Introducing v1.0 of the AI Risk and Reliability Benchmark from MLCommons
- VITAL: A New Dataset for Benchmarking Pluralistic Alignment in Healthcare
- IHEval: Evaluating Language Models on Following the Instruction Hierarchy
- Position: Evaluating Generative AI Systems Is a Social Science Measurement Challenge
- C2SaferRust: Transforming C Projects into Safer Rust with NeuroSymbolic Techniques
- On The Origin of Cultural Biases in Language Models: From Pre-training Data to Linguistic Phenomena
- SafeWatch: An Efficient Safety-Policy Following Video Guardrail Model with Transparent Explanations
- Unified Triplet-Level Hallucination Evaluation for Large Vision-Language Models
- Protecting Privacy in Multimodal Large Language Models with MLLMU-Bench
- SG-Bench: Evaluating LLM Safety Generalization Across Diverse Tasks and Prompt Types
- PAPILLON: Privacy Preservation from Internet-based and Local Language Model Ensembles
- A Comprehensive Evaluation of Cognitive Biases in LLMs
- Can MLLMs Understand the Deep Implication Behind Chinese Images?
- WorldCuisines: A Massive-Scale Benchmark for Multilingual and Multicultural Visual Question Answering on Global Cuisines
- Gender Bias in Decision-Making with Large Language Models: A Study of Relationship Conflicts
- Evaluating Gender Bias of LLMs in Making Morality Judgements
ModSCAN: Measuring Stereotypical Bias in Large Vision-Language Models from Vision and Language Modalities- TLDR: Token-Level Detective Reward Model for Large Vision Language Models
- LabellessFace: Fair Metric Learning for Face Recognition without Attribute Labels
- Testing and Evaluation of Large Language Models: Correctness, Non-Toxicity, and Fairness
- AAVENUE: Detecting LLM Biases on NLU Tasks in AAVE via a Novel Benchmark
- BackdoorLLM: A Comprehensive Benchmark for Backdoor Attacks and Defenses on Large Language Models
- Athena: Safe Autonomous Agents with Verbal Contrastive Learning
- WalledEval: A Comprehensive Safety Evaluation Toolkit for Large Language Models
- Safetywashing: Do AI Safety Benchmarks Actually Measure Safety Progress?
- PERSONA: A Reproducible Testbed for Pluralistic Alignment
- Breaking the Global North Stereotype: A Global South-centric Benchmark Dataset for Auditing and Mitigating Biases in Facial Recognition Systems
- LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models
- AIR-Bench 2024: A Safety Benchmark Based on Risk Categories from Regulations and Policies
- Position: Measure Dataset Diversity, Don't Just Claim It
- Me, Myself, and AI: The Situational Awareness Dataset (SAD) for LLMs
- Social Bias in Large Language Models For Bangla: An Empirical Study on Gender and Religious Bias
- SeqAR: Jailbreak LLMs with Sequential Auto-Generated Characters
- The Art of Saying No: Contextual Noncompliance in Language Models
- MMEvalPro: Calibrating Multimodal Benchmarks Towards Trustworthy and Efficient Evaluation
- Towards Massive Multilingual Holistic Bias
- WildTeaming at Scale: From In-the-Wild Jailbreaks to (Adversarially) Safer Language Models
- The Multilingual Alignment Prism: Aligning Global and Local Preferences to Reduce Harm
- Jailbreaking LLMs with Arabic Transliteration and Arabizi
- CoSafe: Evaluating Large Language Model Safety in Multi-Turn Dialogue Coreference
- Is your benchmark truly adversarial? AdvScore: Evaluating Human-Grounded Adversarialness
- SORRY-Bench: Systematically Evaluating Large Language Model Safety Refusal
- CHiSafetyBench: A Chinese Hierarchical Safety Benchmark for Large Language Models
- OR-Bench: An Over-Refusal Benchmark for Large Language Models
- Safe Multi-agent Reinforcement Learning with Natural Language Constraints
- S-Eval: Towards Automated and Comprehensive Safety Evaluation for Large Language Models
- Poser: Unmasking Alignment Faking LLMs by Manipulating Their Internals
- WildChat: 1M ChatGPT Interaction Logs in the Wild
- FrenchToxicityPrompts: a Large Benchmark for Evaluating and Mitigating Toxicity in French Texts
- Are Watermarks Bugs for Deepfake Detectors? Rethinking Proactive Forensics
- Introducing v0.5 of the AI Safety Benchmark from MLCommons
- On the Robustness of Language Guidance for Low-Level Vision Tasks: Findings from Depth Estimation
- ALERT: A Comprehensive Benchmark for Assessing Large Language Models' Safety through Red Teaming
- JailBreakV: A Benchmark for Assessing the Robustness of MultiModal Large Language Models against Jailbreak Attacks
- IndiBias: A Benchmark Dataset to Measure Social Biases in Language Models for Indian Context
- JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models
- Scaling Behavior of Machine Translation with Large Language Models under Prompt Injection Attacks
- MedSafetyBench: Evaluating and Improving the Medical Safety of Large Language Models
- InjecAgent: Benchmarking Indirect Prompt Injections in Tool-Integrated Large Language Model Agents
- The WMDP Benchmark: Measuring and Reducing Malicious Use With Unlearning
- Abductive Ego-View Accident Video Understanding for Safe Driving Perception
- Controllable Preference Optimization: Toward Controllable Multi-Objective Alignment
- Towards Explainability and Fairness in Swiss Judgement Prediction: Benchmarking on a Multilingual Dataset
- LLMArena: Assessing Capabilities of Large Language Models in Dynamic Multi-Agent Environments
- HypoTermQA: Hypothetical Terms Dataset for Benchmarking Hallucination Tendency of LLMs
- KorNAT: LLM Alignment Benchmark for Korean Social Values and Common Knowledge
- A Unified Framework and Dataset for Assessing Societal Bias in Vision-Language Models
- CIF-Bench: A Chinese Instruction-Following Benchmark for Evaluating the Generalizability of Large Language Models
- ArtPrompt: ASCII Art-based Jailbreak Attacks against Aligned LLMs
- Language Models are Homer Simpson! Safety Re-Alignment of Fine-tuned Language Models through Task Arithmetic
- Benchmark Self-Evolving: A Multi-Agent Framework for Dynamic LLM Evaluation
- A StrongREJECT for Empty Jailbreaks
- SALAD-Bench: A Hierarchical and Comprehensive Safety Benchmark for Large Language Models
- HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal
- Unified Hallucination Detection for Multimodal Large Language Models
- Navigating the OverKill in Large Language Models
- R-Judge: Benchmarking Safety Risk Awareness for LLM Agents
- RAGTruth: A Hallucination Corpus for Developing Trustworthy Retrieval-Augmented Language Models
- Evaluating and Mitigating Discrimination in Language Model Decisions
- Women Are Beautiful, Men Are Leaders: Gender Stereotypes in Machine Translation and Language Modeling
- CDEval: A Benchmark for Measuring the Cultural Dimensions of Large Language Models
- Social Bias Probing: Fairness Benchmarking for Language Models
- SimpleSafetyTests: a Test Suite for Identifying Critical Safety Risks in Large Language Models
- AART: AI-Assisted Red-Teaming with Diverse Data Generation for New LLM-powered Applications
- Flames: Benchmarking Value Alignment of LLMs in Chinese
- PrivLM-Bench: A Multi-level Privacy Evaluation Benchmark for Language Models
- Unveiling Safety Vulnerabilities of Large Language Models
- Can LLMs Follow Simple Rules?
- Tensor Trust: Interpretable Prompt Injection Attacks from an Online Game
- JADE: A Linguistics-based Safety Evaluation Platform for Large Language Models
- Can LLMs Keep a Secret? Testing Privacy Implications of Language Models via Contextual Integrity Theory
- DELPHI: Data for Evaluating LLMs' Performance in Handling Controversial Issues
- Likelihood-based Out-of-Distribution Detection with Denoising Diffusion Probabilistic Models
- Catastrophic Jailbreak of Open-source LLMs via Exploiting Generation
- SWE-bench: Can Language Models Resolve Real-World GitHub Issues?
- Are Personalized Stochastic Parrots More Dangerous? Evaluating Persona Biases in Dialogue Systems
- Can Large Language Models Provide Security & Privacy Advice? Measuring the Ability of LLMs to Refute Misconceptions
- Goal-Oriented Prompt Attack and Safety Evaluation for LLMs
- GPTFUZZER: Red Teaming Large Language Models with Auto-Generated Jailbreak Prompts
- Safety-Tuned LLaMAs: Lessons From Improving the Safety of Large Language Models that Follow Instructions
- FACET: Fairness in Computer Vision Evaluation Benchmark
- CALM : A Multi-task Benchmark for Comprehensive Assessment of Language Model Bias
- Red-Teaming Large Language Models using Chain of Utterances for Safety-Alignment
- Evaluating the Instruction-Following Robustness of Large Language Models to Prompt Injection
- Benchmarking Algorithmic Bias in Face Recognition: An Experimental Approach Using Synthetic Faces and Human Evaluation
- "Do Anything Now": Characterizing and Evaluating In-The-Wild Jailbreak Prompts on Large Language Models
- XSTest: A Test Suite for Identifying Exaggerated Safety Behaviours in Large Language Models
- KoBBQ: Korean Bias Benchmark for Question Answering
- Med-HALT: Medical Domain Hallucination Test for Large Language Models
- Code of "Sirens' Whisper: Inaudible Near-Ultrasonic Jailbreaks of Speech-Driven LLMs"
- Latent Jailbreak: A Benchmark for Evaluating Text Safety and Output Robustness of Large Language Models
- BeaverTails: Towards Improved Safety Alignment of LLM via a Human-Preference Dataset
- On Evaluating and Mitigating Gender Biases in Multilingual Settings
- CBBQ: A Chinese Bias Benchmark Dataset Curated with Human-AI Collaboration for Large Language Models
- WinoQueer: A Community-in-the-Loop Benchmark for Anti-LGBTQ+ Bias in Large Language Models
- DICES Dataset: Diversity in Conversational AI Evaluation for Safety
- CHBias: Bias Evaluation and Mitigation of Chinese Conversational Language Models
- Safety Assessment of Chinese Large Language Models
- HRS-Bench: Holistic, Reliable and Scalable Benchmark for Text-to-Image Models
- Do the Rewards Justify the Means? Measuring Trade-Offs Between Rewards and Ethical Behavior in the MACHIAVELLI Benchmark
- Robo3D: Towards Robust and Reliable 3D Perception against Corruptions
- FVQA 2.0: Introducing Adversarial Samples into Fact-based Visual Question Answering
- MultiRobustBench: Benchmarking Robustness Against Multiple Attacks
- Causally Testing Gender Bias in LLMs: A Case Study on Occupational Bias
- PLUE: Language Understanding Evaluation Benchmark for Privacy Policies in English
- Discovering Language Model Behaviors with Model-Written Evaluations
- On Second Thought, Let's Not Think Step by Step! Bias and Toxicity in Zero-Shot Reasoning
- REAP: A Large-Scale Realistic Adversarial Patch Benchmark
- SafeText: A Benchmark for Exploring Physical Safety in Language Models
- MEDFAIR: Benchmarking Fairness for Medical Imaging
- Re-contextualizing Fairness in NLP: The Case of India
- Red Teaming Language Models to Reduce Harms: Methods, Scaling Behaviors, and Lessons Learned
- SafeBench: A Benchmarking Platform for Safety Evaluation of Autonomous Vehicles
- ProsocialDialog: A Prosocial Backbone for Conversational Agents
- Are Large Pre-Trained Language Models Leaking Your Personal Information?
- "I'm sorry to hear that": Finding New Biases in Language Models with a Holistic Descriptor Dataset
- An Examination of Bias of Facial Analysis based BMI Prediction Models
- Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
- VALUE: Understanding Dialect Disparity in NLU
- ToxiGen: A Large-Scale Machine-Generated Dataset for Adversarial and Implicit Hate Speech Detection
- FairLex: A Multilingual Benchmark for Evaluating Fairness in Legal Text Processing
- Dual use of artificial-intelligence-powered drug discovery
- System Safety and Artificial Intelligence
- The Text Anonymization Benchmark (TAB): A Dedicated Corpus and Evaluation Framework for Text Anonymization
- The Text Anonymization Benchmark (TAB): A Dedicated Corpus and Evaluation Framework for Text Anonymization
- AI and the Everything in the Whole Wide World Benchmark
- Adversarial GLUE: A Multi-Task Benchmark for Robustness Evaluation of Language Models
- Attention-like processes in the Drosophila brain
- On the Safety of Conversational Models: Taxonomy, Dataset, and Benchmark
- BBQ: A Hand-Built Bias Benchmark for Question Answering
- BBQ: A hand-built bias benchmark for question answering
- ePiC: Employing Proverbs in Context as a Benchmark for Abstract Language Understanding
- Mitigating Language-Dependent Ethnic Bias in BERT
- Towards Understanding and Mitigating Social Biases in Language Models
- Understanding and Evaluating Racial Biases in Image Captioning
- RedditBias: A Real-World Resource for Bias Evaluation and Debiasing of Conversational Language Models
- Adversarial VQA: A New Benchmark for Evaluating the Robustness of VQA Models
- Patterns, predictions, and actions: A story about machine learning
- Bias Out-of-the-Box: An Empirical Analysis of Intersectional Occupational Biases in Popular Generative Language Models
- BOLD: Dataset and Metrics for Measuring Biases in Open-Ended Language Generation
- Re-imagining Algorithmic Fairness in India and Beyond
- RADDLE: An Evaluation Benchmark and Analysis Platform for Robust Task-oriented Dialog Systems
- Extracting Training Data from Large Language Models
- WILDS: A Benchmark of in-the-Wild Distribution Shifts
- From Hero to Zéroe: A Benchmark of Low-Level Adversarial Attacks
- UnQovering Stereotyping Biases via Underspecified Questions
- CrowS-Pairs: A Challenge Dataset for Measuring Social Biases in Masked Language Models
- RealToxicityPrompts: Evaluating Neural Toxic Degeneration in Language Models
- On the Threat of npm Vulnerable Dependencies in Node.js Applications
- Can Autonomous Vehicles Identify, Recover From, and Adapt to Distribution Shifts?
- Roses Are Red, Violets Are Blue... but Should Vqa Expect Them To?
- Beyond Accuracy: Behavioral Testing of NLP models with CheckList
- StereoSet: Measuring stereotypical bias in pretrained language models
- Measurement and Fairness
- Does Gender Matter? Towards Fairness in Dialogue Systems
- MLPerf Training Benchmark
- The Woman Worked as a Babysitter: On Biases in Language Generation
- Build it Break it Fix it for Dialogue Safety: Robustness from Adversarial Human Attack
- BenchPress: Analyzing Android App Vulnerability Benchmark Suites
- Fairness and Abstraction in Sociotechnical Systems
- Examining Gender and Race Bias in Two Hundred Sentiment Analysis Systems
- Gender Bias in Coreference Resolution
- Gender Bias in Coreference Resolution: Evaluation and Debiasing Methods
- The Malicious Use of Artificial Intelligence: Forecasting, Prevention, and Mitigation
- Children and the Data Cycle: Rights and Ethics in a Big Data World
- Saturation in qualitative research: exploring its conceptualization and operationalization
- Achieving non-discrimination in prediction
- Concrete Problems in AI Safety
- ImageNet Large Scale Visual Recognition Challenge
- ImageNet Large Scale Visual Recognition Challenge
- Assessing the impact of planned social change
- On The Quantitative Definition of Risk
- TiEBe: Tracking Language Model Recall of Notable Worldwide Events Through Time
Cited by