Large Language Models Are Not Robust Multiple Choice Selectors
2023/09/07 by Chujie Zheng, Hao Zhou, Zheng, Chujie +7 · 1 voice · 117 citations
Computer Science · #Expert finding and Q&A systems #Natural Language Processing Techniques #Topic Modeling #cs.CL
paper · pdf · doi:10.48550/arxiv.2309.03882
openalex publication_date 2023/09/07 · openalex created_date 2025/10/10 · openalex updated_date 2026/08/04
Abstract
Multiple choice questions (MCQs) serve as a common yet important task format in the evaluation of large language models (LLMs). This work shows that modern LLMs are vulnerable to option position changes in MCQs due to their inherent "selection bias", namely, they prefer to select specific option IDs as answers (like "Option A"). Through extensive empirical analyses with 20 LLMs on three benchmarks, we pinpoint that this behavioral bias primarily stems from LLMs' token bias, where the model a priori assigns more probabilistic mass to specific option ID tokens (e.g., A/B/C/D) when predicting answers from the option IDs. To mitigate selection bias, we propose a label-free, inference-time debiasing method, called PriDe, which separates the model's prior bias for option IDs from the overall prediction distribution. PriDe first estimates the prior by permutating option contents on a small number of test samples, and then applies the estimated prior to debias the remaining samples. We demonstrate that it achieves interpretable and transferable debiasing with high computational efficiency. We hope this work can draw broader research attention to the bias and robustness of modern LLMs.
Cited by
- GamiBench: Evaluating Spatial Reasoning and 2D-to-3D Planning Capabilities of MLLMs with Origami Folding Tasks
- VeraGrid-Agent: Tool-Augmented LLMs for Distribution Optimal Power Flow at the Grid Edge
- LLM-Ideoplasticity: Measuring Ideological Plasticity in the Political Behavior of LLMs as a Context-Conditioned Distribution
- Measuring and Improving Behavioral Consistency in Large Language Models through Fact-Heuristic-Emotion State Enforcement
- Embodied4C: Measuring What Matters for Embodied Vision-Language Navigation
- TriDF: Evaluating Perception, Detection, and Hallucination for Interpretable DeepFake Detection
- CoSPlan: Corrective Sequential Planning via Scene Graph Incremental Updates
- Step-by-step Layered Design Generation
- Improving Score Reliability of Multiple Choice Benchmarks with Consistency Evaluation and Altered Answer Choices
- Unexplored flaws in multiple-choice VQA evaluations
- Building Domain-Specific Small Language Models via Guided Data Generation
- Beyond Multiple Choice: Verifiable OpenQA for Robust Vision-Language RFT
- HEAD-QA v2: Expanding a Healthcare Benchmark for Reasoning
- Quantifying and Mitigating Selection Bias in LLMs: A Transferable LoRA Fine-Tuning and Efficient Majority Voting Approach
- Unifying points of interest taxonomies: mapping OpenStreetMap tags to the Foursquare category system
- A Multifaceted Analysis of Negative Bias in Large Language Models through the Lens of Parametric Knowledge
- What's in Common? Multimodal Models Hallucinate When Reasoning Across Scenes
- Lost in Phonation: Voice Quality Variation as an Evaluation Dimension for Speech Foundation Models
- Stemma: Induced Decision Regions Reveal LLM Provenance
- MultivationBench: A Benchmark for Multimodal Sequential Motivation Reasoning
- When LLMs Agree, Are They Right? Auditing Self-Consistency and Cross-Model Agreement as Confidence Signals
- Finding Culture-Sensitive Neurons in Vision-Language Models
- Beyond Understanding: Evaluating the Pragmatic Gap in LLMs' Cultural Processing of Figurative Language
- VISTA: A Test-Time Self-Improving Video Generation Agent
- Investigating the Impact of Rationales for LLMs on Natural Language Understanding
- Selective Adversarial Attacks on LLM Benchmarks
- Multi-Agent Debate for LLM Judges with Adaptive Stability Detection
- Too Open for Opinion? Embracing Open-Endedness in Large Language Models for Social Simulation
- GOLD PANNING: Strategic Context Shuffling for Needle-in-Haystack Reasoning
- MDSEval: A Meta-Evaluation Benchmark for Multimodal Dialogue Summarization
- If Probable, Then Acceptable? Understanding Conditional Acceptability Judgments in Large Language Models
- Reasoning by Exploration: A Unified Approach to Retrieval and Generation over Graphs
- MMA-ASIA: A Multilingual and Multimodal Alignment Framework for Culturally-Grounded Evaluation
- FinReflectKG -- EvalBench: Benchmarking Financial KG with Multi-Dimensional Evaluation
- Robustness assessment of large audio language models in multiple-choice evaluation
- MHA-RAG: Improving Efficiency, Accuracy, and Consistency by Encoding Exemplars as Soft Prompts
- IndiCASA: A Dataset and Bias Evaluation Framework in LLMs Using Contrastive Embedding Similarity in the Indian Context
- When Voice Matters: Evidence of Gender Disparity in Positional Bias of SpeechLLMs
- Hearing the Order: Investigating Selection Bias in Large Audio-Language Models
- Are Large Language Models Chronically Online Surfers? A Dataset for Chinese Internet Meme Explanation
- MEMTRACK: Evaluating Long-Term Memory and State Tracking in Multi-Platform Dynamic Agent Environments
- BiasBusters: Uncovering and Mitigating Tool Selection Bias in Large Language Models
- Bias Mitigation or Cultural Commonsense? Evaluating LLMs with a Japanese Dataset
- Query Circuits: Explaining How Language Models Answer User Prompts
- Bridging the Knowledge-Prediction Gap in LLMs on Multiple-Choice Questions
- JGU Mainz's Submission to the WMT25 Shared Task on LLMs with Limited Resources for Slavic Languages: MT and QA
- Debiasing Large Language Models in Thai Political Stance Detection via Counterfactual Calibration
- Instruction Boundary: Quantifying Biases in LLM Reasoning under Various Coverage
- What Does Your Benchmark Really Measure? A Framework for Robust Inference of AI Capabilities
- SMITE: Enhancing Fairness in LLMs through Optimal In-Context Example Selection via Dynamic Validation
- The Ranking Blind Spot: Decision Hijacking in LLM-based Text Ranking
- Benchmarking and Mitigating MCQA Selection Bias of Large Vision-Language Models
- Debate or Vote: Which Yields Better Decisions in Multi-Agent Large Language Models?
- MORABLES: A Benchmark for Assessing Abstract Moral Reasoning in LLMs with Fables
- Maestro: Self-Improving Text-to-Image Generation via Agent Orchestration
- Are Humans as Brittle as Large Language Models?
- Multi-view-guided Passage Reranking with Large Language Models
- When Models Refuse: Political Steerability and Feature Richness as Measures of Ideological Depth
- INSEva: A Comprehensive Chinese Benchmark for Large Language Models in Insurance
- CrossHOI-Bench: A Unified Benchmark for HOI Evaluation across Vision-Language Models and HOI-Specific Methods
- VMMU: A Vietnamese Multitask Multimodal Understanding and Reasoning Benchmark
- Jointly Generating and Attributing Answers using Logits of Document-Identifier Tokens
- SKATE, a Scalable Tournament Eval: Weaker LLMs differentiate between stronger ones using verifiable challenges
- Investigating Hallucination in Conversations for Low Resource Languages
- Metric assessment protocol in the context of answer fluctuation on MCQ tasks
- AgriEval: A Comprehensive Chinese Agricultural Benchmark for Large Language Models
- Reasoning Models are Test Exploiters: Rethinking Multiple-Choice
- AlgoSimBench: Identifying Algorithmically Similar Problems for Competitive Programming
- SCOPE: Stochastic and Counterbiased Option Placement for Evaluating Large Language Models
- Large Language Models as Medical Codes Selectors: a benchmark using the International Classification of Primary Care
- Exploiting Primacy Effect To Improve Large Language Models
- Analysis of Threat-Based Manipulation in Large Language Models: A Dual Perspective on Vulnerabilities and Performance Enhancement Opportunities
- Do LLMs Give Psychometrically Plausible Responses in Educational Assessments?
- The Generative Energy Arena (GEA): Incorporating Energy Awareness in Large Language Model (LLM) Human Evaluations
- Revisiting LLM Value Probing Strategies: Are They Robust and Expressive?
- A Training-Free, Task-Agnostic Framework for Enhancing MLLM Performance on High-Resolution Images
- XiYan-SQL: A Novel Multi-Generator Framework For Text-to-SQL
- Answer Matching Outperforms Multiple Choice for Language Model Evaluation
- MI-CXR: A Benchmark for Longitudinal Reasoning over Multi-Interval Chest X-rays
- Quantifying Cognitive Bias Induction in LLM-Generated Content
- EvalAssist: A Human-Centered Tool for LLM-as-a-Judge
- Positional Bias in Binary Question Answering: How Uncertainty Shapes Model Preferences
- SoMi-ToM: Evaluating Multi-Perspective Theory of Mind in Embodied Social Interactions
- Agent-to-Agent Theory of Mind: Testing Interlocutor Awareness among Large Language Models
- Lost at the Beginning of Reasoning
- From Calibration to Collaboration: LLM Uncertainty Quantification Should Be More Human-Centered
- Benchmarking the Pedagogical Knowledge of Large Language Models
- Gazal-R1: Achieving State-of-the-Art Medical Reasoning with Parameter-Efficient Two-Stage Training
- The Effect of State Representation on LLM Agent Behavior in Dynamic Routing Games
- MotiveBench: How Far Are We From Human-Like Motivational Reasoning in Large Language Models?
- Burn After Reading: Do Multimodal Large Language Models Truly Capture Order of Events in Image Sequences?
- Quantifying Cross-Modality Memorization in Vision-Language Models
- Evaluation is All You Need: Strategic Overclaiming of LLM Reasoning Capabilities Through Evaluation Design
- FailureSensorIQ: A Multi-Choice QA Dataset for Understanding Sensor Relationships and Failure Modes
- Existing Large Language Model Unlearning Evaluations Are Inconclusive
- SATA-BENCH: Select All That Apply Benchmark for Multiple Choice Questions
- ClinBench-HPB: A Clinical Benchmark for Evaluating LLMs in Hepato-Pancreato-Biliary Diseases
- Do Large Language Models Think Like the Brain? Sentence-Level Evidences from Layer-Wise Embeddings and fMRI
- Hidden in the Haystack: Smaller Needles are More Difficult for LLMs to Find
- Let Androids Dream of Electric Sheep: A Human-Inspired Image Implication Understanding and Reasoning Framework
- What Media Frames Reveal About Stance: A Dataset and Study about Memes in Climate Change Discourse
- QA-prompting: Improving Summarization with Large Language Models using Question-Answering
- HumMusQA: A Human-written Music Understanding QA Benchmark Dataset
- Same Meaning, Different Scores: Lexical and Syntactic Sensitivity in LLM Evaluation
- Systematic Bias in Large Language Models: Discrepant Response Patterns in Binary vs. Continuous Judgment Tasks
- KETCHUP: K-Step Return Estimation for Sequential Knowledge Distillation
- When2Call: When (not) to Call Tools
- AI Achieves a Perfect LSAT Score
- SciHorizon-GENE: Benchmarking LLM for Life Sciences Inference from Gene Knowledge to Functional Understanding
- When Models Decide and When They Bind: A Two-Stage Computation for Multiple-Choice Question-Answering
- Every Wrong Answer Counts: Option-Level Psychometrics for LLM Multiple-Choice Benchmarks
- IslamicTurathBench: A Multi-Task, Multi-Discipline Benchmark for Evaluating Large Language Models on the Islamic Scholarly Tradition (turath)
- The Nuclear Decision-Making Benchmark: Evaluating Frontier LLMs on Nuclear Tendencies
- Improving Instruct Models for Free: A Study on Partial Adaptation
- CameraBench: Benchmarking Visual Reasoning in MLLMs via Photography
- Large Language Models Could Be Rote Learners
- Beyond Self-Reports: Multi-Observer Agents for Personality Assessment in Large Language Models
Discussions
Related