Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators
2024/04/06 by Yann Dubois, Dubois, Yann, Balázs Galambosi +5 · 127 citations
Computer Science · #AI in Service Interactions #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Natural Language Processing Techniques #Topic Modeling
paper · pdf · doi:10.48550/arxiv.2404.04475
openalex publication_date 2024/04/06 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Abstract
LLM-based auto-annotators have become a key component of the LLM development process due to their cost-effectiveness and scalability compared to human-based evaluation. However, these auto-annotators can introduce biases that are hard to remove. Even simple, known confounders such as preference for longer outputs remain in existing automated evaluation metrics. We propose a simple regression analysis approach for controlling biases in auto-evaluations. As a real case study, we focus on reducing the length bias of AlpacaEval, a fast and affordable benchmark for instruction-tuned LLMs that uses LLMs to estimate response quality. Despite being highly correlated with human preferences, AlpacaEval is known to favor models that generate longer outputs. We introduce a length-controlled AlpacaEval that aims to answer the counterfactual question: "What would the preference be if the model's and baseline's output had the same length?" To achieve this, we first fit a generalized linear model to predict the biased auto-annotator's preferences based on the mediators we want to control for (length difference) and other relevant features. We then obtain length-controlled preferences by predicting preferences while conditioning the GLM with a zero difference in lengths. Length-controlling not only improves the robustness of the metric to manipulations in model verbosity, but we also find that it increases the Spearman correlation with LMSYS Chatbot Arena from 0.94 to 0.98.
Cited by
- Eliminating Inductive Bias in Reward Models with Information-Theoretic Guidance
- Accuracy Hides How Language Models Fail: Measuring Failure States Under Matched Output Budgets
- Less Data, Better Alignment: Data-Centric Multi-Evaluator Agreement for Preference Optimization
- Inverse RL Helps Align AI by Imitating Humans
- Safety Alignment of LMs via Non-cooperative Games
- AWPO: Enhancing Tool-Use of Large Language Models through Adaptive Integration of Reasoning Rewards
- Recontextualization Mitigates Specification Gaming without Modifying the Specification
- Stackelberg Learning from Human Feedback: Preference Optimization as a Sequential Game
- Revisiting the Reliability of Language Models in Instruction-Following
- Causal Judge Evaluation: Calibrated Surrogate Metrics for LLM Systems
- ARM-Thinker: Reinforcing Multimodal Generative Reward Models with Agentic Tool Use and Visual Reasoning
- Attention is All You Need to Defend Against Indirect Prompt Injection Attacks in LLMs
- Beyond Prototyping: Autonomous, Enterprise-Grade Frontend Development from Pixel to Production via a Specialized Multi-Agent Framework
- What Is Preference Optimization Doing, How and Why?
- Ambiguity Awareness Optimization: Towards Semantic Disambiguation for Direct Preference Optimization
- Experts are all you need: A Composable Framework for Large Language Model Inference
- On Evaluating LLM Alignment by Evaluating LLMs as Judges
- ABM-LoRA: Activation Boundary Matching for Fast Convergence in Low-Rank Adaptation
- Cognitive Foundations for Reasoning and Their Manifestation in LLMs
- Plan-and-Write: Structure-Guided Length Control for LLMs without Model Retraining
- VIDEOP2R: Video Understanding from Perception to Reasoning
- Steering Pretrained Drafters during Speculative Decoding
- Beyond Task-Oriented and Chitchat Dialogues: Proactive and Transition-Aware Conversational Agents
- SERL: Self-Examining Reinforcement Learning on Open-Domain
- CoLM: Collaborative Large Models via A Client-Server Paradigm
- Textual Self-attention Network: Test-Time Preference Optimization through Textual Gradient-based Attention
- No-Human in the Loop: Agentic Evaluation at Scale for Recommendation
- PublicAgent: Multi-Agent Design Principles From an LLM-Based Open Data Analysis Framework
- LLM-as-a-Verifier: A General-Purpose Verification Framework
- Meta-Learned Reward Shaping for Reinforcement Learning from Human Feedback
- Mediocrity is the key for LLM as a Judge Anchor Selection
- Ideas Have Genomes: Benchmarking Scientific Lineage Reasoning and Lineage-Grounded Idea Generation
- PRISMM-Bench: A Benchmark of Peer-Review Grounded Multimodal Inconsistencies
- Judging Against the Reference: Uncovering Knowledge-Driven Failures in LLM-Judges on QA Evaluation
- MoReBench: Evaluating Procedural and Pluralistic Moral Reasoning in Language Models, More than Outcomes
- Semi-Supervised Preference Optimization with Limited Feedback
- SelecTKD: Selective Token-Weighted Knowledge Distillation for LLMs
- GIFT: Group-relative Implicit Fine Tuning Integrates GRPO with DPO and UNA
- Offline Preference Optimization via Maximum Marginal Likelihood Estimation
- Feature-Guided SAE Steering for Refusal-Rate Control using Contrasting Prompts
- Beyond Pairwise: Empowering LLM Alignment With Ranked Choice Modeling
- Self-Rewarding PPO: Aligning Large Language Models with Demonstrations Only
- Language Ranker: A Lightweight Ranking framework for LLM Decoding
- Ask a Strong LLM Judge when Your Reward Model is Uncertain
- A Principle-based Framework for the Development and Evaluation of Large Language Models for Health and Wellness
- Context-level Language Modeling by Learning Predictive Context Embeddings
- Data-Centric Lessons To Improve Speech-Language Pretraining
- Defending Against Prompt Injection with DataFilter
- Rectifying Shortcut Behaviors in Preference-based Reward Learning
- WebDevJudge: Evaluating (M)LLMs as Critiques for Web Development Quality
- ChronoPlay: A Framework for Modeling Dual Dynamics and Authenticity in Game RAG Benchmarks
- Towards Faithful and Controllable Personalization via Critique-Post-Edit Reinforcement Learning
- Planned Diffusion
- Foundational Automatic Evaluators: Scaling Multi-Task Generative Evaluator Training for Reasoning-Centric Domains
- PANER: A Paraphrase-Augmented Framework for Low-Resource Named Entity Recognition
- The Atomic Instruction Gap: Instruction-Tuned LLMs Struggle with Simple, Self-Contained Directives
- Auto-Rubric: Learning From Implicit Weights to Explicit Rubrics for Reward Modeling
- Reasoning with Sampling: Your Base Model is Smarter Than You Think
- RLSR: Reinforcement Learning with Supervised Reward Outperforms SFT in Instruction Following
- DSCD: Large Language Model Detoxification with Self-Constrained Decoding
- On the Role of Preference Variance in Preference Optimization
- Faster LLM Inference via Sequential Monte Carlo
- IatroBench: Pre-Registered Evidence of Iatrogenic Harm from AI Safety Measures
- Safety Game: Balancing Safe and Informative Conversations with Blackbox Agentic AI using LP Solvers
- Alif: Advancing Urdu Large Language Models via Multilingual Synthetic Data Distillation
- RefGrader: Automated Grading of Mathematical Competition Proofs using Agentic Workflows
- The Alignment Waltz: Jointly Training Agents to Collaborate for Safety
- Metric Calculating Benchmark: Code-Verifiable Complicate Instruction Following Benchmark for Large Language Models
- Contrastive Weak-to-strong Generalization
- PIKA: Expert-Level Synthetic Datasets for Post-Training Alignment from Scratch
- Online Rubrics Elicitation from Pairwise Comparisons
- Provably Mitigating Corruption, Overoptimization, and Verbosity Simultaneously in Offline and Online RLHF/DPO Alignment
- Staircase Streaming for Low-Latency Multi-Agent Inference
- Reward Model Routing in Alignment
- Safety Instincts: LLMs Learn to Trust Their Internal Compass for Self-Defense
- Beyond Log Likelihood: Probability-Based Objectives for Supervised Fine-Tuning across the Model Capability Continuum
- The Silent Judge: Unacknowledged Shortcut Bias in LLM-as-a-Judge
- The Era of Real-World Human Interaction: RL from User Conversations
- CDT: A Comprehensive Capability Framework for Large Language Models Across Cognition, Domain, and Task
- Humanline: Online Alignment as Perceptual Loss
- RE-PO: Robust Enhanced Policy Optimization as a General Framework for LLM Alignment
- Don't Settle Too Early: Self-Reflective Remasking for Diffusion Language Models
- Alignment through Meta-Weighted Online Sampling: Bridging the Gap between Data Generation and Preference Optimization
- p-less Sampling: A Robust Hyperparameter-Free Approach for LLM Decoding
- General Exploratory Bonus for Optimistic Exploration in RLHF
- Effective Quantization of Muon Optimizer States
- Multiplayer Nash Preference Optimization
- Adaptive Margin RLHF via Preference over Preferences
- Semantic Voting: A Self-Evaluation-Free Approach for Efficient LLM Self-Improvement on Unverifiable Open-ended Tasks
- S2J: Bridging the Gap Between Solving and Judging Ability in Generative Reward Models
- PSRT: Accelerating LRM-based Guard Models via Prefilled Safe Reasoning Traces
- Bidirectional Intention Inference Enhances LLMs' Defense Against Multi-Turn Jailbreak Attacks
- TrustJudge: Inconsistencies of LLM-as-a-Judge and How to Alleviate Them
- RLBFF: Binary Flexible Feedback to bridge between Human Feedback & Verifiable Rewards
- BridgeAlign: Bridging Preference Alignment for Humanities and Social Sciences
- HSS-Synth: Humanities and Social Sciences Data Synthesis for LLMs
- Proximal Supervised Fine-Tuning
- Enhancing Speech Large Language Models through Reinforced Behavior Alignment
- Weights-Rotated Preference Optimization for Large Language Models
- Preference Distillation via Value based Reinforcement Learning
- The Anatomy of Alignment: Decomposing Preference Optimization by Steering Sparse Features
- What Matters in Data for DPO?
- Accelerate Scaling of LLM Finetuning via Quantifying the Coverage and Depth of Instruction Set
- GeoArena: An Open Platform for Benchmarking Large Vision-language Models on WorldWide Image Geolocalization
- Adaptive Preference Optimization with Uncertainty-aware Utility Anchor
- Retrieval-Augmented Defense: Adaptive and Controllable Jailbreak Prevention for Large Language Models
- Unlearning That Lasts: Utility-Preserving, Robust, and Almost Irreversible Forgetting in LLMs
- Jointly Reinforcing Diversity and Quality in Language Model Generations
- ResearchQA: Evaluating Scholarly Question Answering at Scale Across 75 Fields with Survey-Mined Questions and Rubrics
- UI-Bench: A Benchmark for Evaluating Design Capabilities of AI Text-to-App Tools
- IntentionReasoner: Facilitating Adaptive LLM Safeguards through Intent Reasoning and Selective Query Refinement
- ReSURE: Regularizing Supervision Unreliability for Multi-turn Dialogue Fine-tuning
- HEAL: A Hypothesis-Based Preference-Aware Analysis Framework
- Generative Interfaces for Language Models
- Rethinking Caching for LLM Serving Systems: Beyond Traditional Heuristics
- DEPTH: Hallucination-Free Relation Extraction via Dependency-Aware Sentence Simplification and Two-tiered Hierarchical Refinement
- The illusion of a perfect metric: Why evaluating AI's words is harder than it looks
- Prompt-Based One-Shot Exact Length-Controlled Generation with LLMs
- Bridging Human and LLM Judgments: Understanding and Narrowing the Gap
- Diversity First, Quality Later: A Two-Stage Assumption for Language Model Alignment
- Exploring Safety Alignment Evaluation of LLMs in Chinese Mental Health Dialogues via LLM-as-Judge
- Data Selection for LLM Alignment Using Fine-Grained Preferences
- Enhancing Small LLM Alignment through Margin-Based Objective Modifications under Resource Constraints
- Context Misleads LLMs: The Role of Context Filtering in Maintaining Safe Alignment of LLMs
- Synthesizing scientific literature with retrieval-augmented language models
- DP-LLM: Runtime Model Adaptation with Dynamic Layer-wise Precision Assignment
- Llama-3.1-FoundationAI-SecurityLLM-8B-Instruct Technical Report
- Cascaded Information Disclosure for Generalized Evaluation of Problem Solving Capabilities
- CoT-Self-Instruct: Building high-quality synthetic prompts for reasoning and non-reasoning tasks
Related