Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators
2024/04/06 by Yann Dubois, Dubois, Yann, Balázs Galambosi +5 · 247 citations
Computer Science · #AI in Service Interactions #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Natural Language Processing Techniques #Topic Modeling
paper · pdf · doi:10.48550/arxiv.2404.04475
openalex publication_date 2024/04/06 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Abstract
LLM-based auto-annotators have become a key component of the LLM development process due to their cost-effectiveness and scalability compared to human-based evaluation. However, these auto-annotators can introduce biases that are hard to remove. Even simple, known confounders such as preference for longer outputs remain in existing automated evaluation metrics. We propose a simple regression analysis approach for controlling biases in auto-evaluations. As a real case study, we focus on reducing the length bias of AlpacaEval, a fast and affordable benchmark for instruction-tuned LLMs that uses LLMs to estimate response quality. Despite being highly correlated with human preferences, AlpacaEval is known to favor models that generate longer outputs. We introduce a length-controlled AlpacaEval that aims to answer the counterfactual question: "What would the preference be if the model's and baseline's output had the same length?" To achieve this, we first fit a generalized linear model to predict the biased auto-annotator's preferences based on the mediators we want to control for (length difference) and other relevant features. We then obtain length-controlled preferences by predicting preferences while conditioning the GLM with a zero difference in lengths. Length-controlling not only improves the robustness of the metric to manipulations in model verbosity, but we also find that it increases the Spearman correlation with LMSYS Chatbot Arena from 0.94 to 0.98.
Cited by
- Towards Bridging the Reward-Generation Gap in Direct Alignment Algorithms
- Eliminating Inductive Bias in Reward Models with Information-Theoretic Guidance
- Accuracy Hides How Language Models Fail: Measuring Failure States Under Matched Output Budgets
- Less Data, Better Alignment: Data-Centric Multi-Evaluator Agreement for Preference Optimization
- Inverse RL Helps Align AI by Imitating Humans
- Safety Alignment of LMs via Non-cooperative Games
- AWPO: Enhancing Tool-Use of Large Language Models through Adaptive Integration of Reasoning Rewards
- Recontextualization Mitigates Specification Gaming without Modifying the Specification
- Stackelberg Learning from Human Feedback: Preference Optimization as a Sequential Game
- Revisiting the Reliability of Language Models in Instruction-Following
- Causal Judge Evaluation: Calibrated Surrogate Metrics for LLM Systems
- ARM-Thinker: Reinforcing Multimodal Generative Reward Models with Agentic Tool Use and Visual Reasoning
- Attention is All You Need to Defend Against Indirect Prompt Injection Attacks in LLMs
- Beyond Prototyping: Autonomous, Enterprise-Grade Frontend Development from Pixel to Production via a Specialized Multi-Agent Framework
- What Is Preference Optimization Doing, How and Why?
- Ambiguity Awareness Optimization: Towards Semantic Disambiguation for Direct Preference Optimization
- Experts are all you need: A Composable Framework for Large Language Model Inference
- On Evaluating LLM Alignment by Evaluating LLMs as Judges
- ABM-LoRA: Activation Boundary Matching for Fast Convergence in Low-Rank Adaptation
- Cognitive Foundations for Reasoning and Their Manifestation in LLMs
- Plan-and-Write: Structure-Guided Length Control for LLMs without Model Retraining
- VIDEOP2R: Video Understanding from Perception to Reasoning
- Steering Pretrained Drafters during Speculative Decoding
- Beyond Task-Oriented and Chitchat Dialogues: Proactive and Transition-Aware Conversational Agents
- SERL: Self-Examining Reinforcement Learning on Open-Domain
- CoLM: Collaborative Large Models via A Client-Server Paradigm
- Textual Self-attention Network: Test-Time Preference Optimization through Textual Gradient-based Attention
- No-Human in the Loop: Agentic Evaluation at Scale for Recommendation
- PublicAgent: Multi-Agent Design Principles From an LLM-Based Open Data Analysis Framework
- LLM-as-a-Verifier: A General-Purpose Verification Framework
- Meta-Learned Reward Shaping for Reinforcement Learning from Human Feedback
- Mediocrity is the key for LLM as a Judge Anchor Selection
- Ideas Have Genomes: Benchmarking Scientific Lineage Reasoning and Lineage-Grounded Idea Generation
- VerIF: Verification Engineering for Reinforcement Learning in Instruction Following
- PRISMM-Bench: A Benchmark of Peer-Review Grounded Multimodal Inconsistencies
- Judging Against the Reference: Uncovering Knowledge-Driven Failures in LLM-Judges on QA Evaluation
- MoReBench: Evaluating Procedural and Pluralistic Moral Reasoning in Language Models, More than Outcomes
- Semi-Supervised Preference Optimization with Limited Feedback
- SelecTKD: Selective Token-Weighted Knowledge Distillation for LLMs
- GIFT: Group-relative Implicit Fine Tuning Integrates GRPO with DPO and UNA
- Offline Preference Optimization via Maximum Marginal Likelihood Estimation
- Feature-Guided SAE Steering for Refusal-Rate Control using Contrasting Prompts
- Beyond Pairwise: Empowering LLM Alignment With Ranked Choice Modeling
- Self-Rewarding PPO: Aligning Large Language Models with Demonstrations Only
- Language Ranker: A Lightweight Ranking framework for LLM Decoding
- Ask a Strong LLM Judge when Your Reward Model is Uncertain
- A Principle-based Framework for the Development and Evaluation of Large Language Models for Health and Wellness
- Context-level Language Modeling by Learning Predictive Context Embeddings
- Data-Centric Lessons To Improve Speech-Language Pretraining
- Defending Against Prompt Injection with DataFilter
- Rectifying Shortcut Behaviors in Preference-based Reward Learning
- WebDevJudge: Evaluating (M)LLMs as Critiques for Web Development Quality
- ChronoPlay: A Framework for Modeling Dual Dynamics and Authenticity in Game RAG Benchmarks
- Towards Faithful and Controllable Personalization via Critique-Post-Edit Reinforcement Learning
- Planned Diffusion
- Foundational Automatic Evaluators: Scaling Multi-Task Generative Evaluator Training for Reasoning-Centric Domains
- PANER: A Paraphrase-Augmented Framework for Low-Resource Named Entity Recognition
- The Atomic Instruction Gap: Instruction-Tuned LLMs Struggle with Simple, Self-Contained Directives
- Auto-Rubric: Learning From Implicit Weights to Explicit Rubrics for Reward Modeling
- Reasoning with Sampling: Your Base Model is Smarter Than You Think
- RLSR: Reinforcement Learning with Supervised Reward Outperforms SFT in Instruction Following
- DSCD: Large Language Model Detoxification with Self-Constrained Decoding
- On the Role of Preference Variance in Preference Optimization
- Faster LLM Inference via Sequential Monte Carlo
- IatroBench: Pre-Registered Evidence of Iatrogenic Harm from AI Safety Measures
- Safety Game: Balancing Safe and Informative Conversations with Blackbox Agentic AI using LP Solvers
- Alif: Advancing Urdu Large Language Models via Multilingual Synthetic Data Distillation
- RefGrader: Automated Grading of Mathematical Competition Proofs using Agentic Workflows
- The Alignment Waltz: Jointly Training Agents to Collaborate for Safety
- Metric Calculating Benchmark: Code-Verifiable Complicate Instruction Following Benchmark for Large Language Models
- Contrastive Weak-to-strong Generalization
- PIKA: Expert-Level Synthetic Datasets for Post-Training Alignment from Scratch
- Online Rubrics Elicitation from Pairwise Comparisons
- Provably Mitigating Corruption, Overoptimization, and Verbosity Simultaneously in Offline and Online RLHF/DPO Alignment
- Staircase Streaming for Low-Latency Multi-Agent Inference
- Reward Model Routing in Alignment
- Safety Instincts: LLMs Learn to Trust Their Internal Compass for Self-Defense
- Beyond Log Likelihood: Probability-Based Objectives for Supervised Fine-Tuning across the Model Capability Continuum
- The Silent Judge: Unacknowledged Shortcut Bias in LLM-as-a-Judge
- The Era of Real-World Human Interaction: RL from User Conversations
- CDT: A Comprehensive Capability Framework for Large Language Models Across Cognition, Domain, and Task
- Humanline: Online Alignment as Perceptual Loss
- RE-PO: Robust Enhanced Policy Optimization as a General Framework for LLM Alignment
- Don't Settle Too Early: Self-Reflective Remasking for Diffusion Language Models
- Alignment through Meta-Weighted Online Sampling: Bridging the Gap between Data Generation and Preference Optimization
- p-less Sampling: A Robust Hyperparameter-Free Approach for LLM Decoding
- General Exploratory Bonus for Optimistic Exploration in RLHF
- Effective Quantization of Muon Optimizer States
- Multiplayer Nash Preference Optimization
- Adaptive Margin RLHF via Preference over Preferences
- Semantic Voting: A Self-Evaluation-Free Approach for Efficient LLM Self-Improvement on Unverifiable Open-ended Tasks
- S2J: Bridging the Gap Between Solving and Judging Ability in Generative Reward Models
- PSRT: Accelerating LRM-based Guard Models via Prefilled Safe Reasoning Traces
- Bidirectional Intention Inference Enhances LLMs' Defense Against Multi-Turn Jailbreak Attacks
- TrustJudge: Inconsistencies of LLM-as-a-Judge and How to Alleviate Them
- RLBFF: Binary Flexible Feedback to bridge between Human Feedback & Verifiable Rewards
- BridgeAlign: Bridging Preference Alignment for Humanities and Social Sciences
- HSS-Synth: Humanities and Social Sciences Data Synthesis for LLMs
- Proximal Supervised Fine-Tuning
- Enhancing Speech Large Language Models through Reinforced Behavior Alignment
- Weights-Rotated Preference Optimization for Large Language Models
- Preference Distillation via Value based Reinforcement Learning
- The Anatomy of Alignment: Decomposing Preference Optimization by Steering Sparse Features
- What Matters in Data for DPO?
- Accelerate Scaling of LLM Finetuning via Quantifying the Coverage and Depth of Instruction Set
- GeoArena: Evaluating Open-World Geographic Reasoning in Large Vision-Language Models
- Adaptive Preference Optimization with Uncertainty-aware Utility Anchor
- Retrieval-Augmented Defense: Adaptive and Controllable Jailbreak Prevention for Large Language Models
- Unlearning That Lasts: Utility-Preserving, Robust, and Almost Irreversible Forgetting in LLMs
- Jointly Reinforcing Diversity and Quality in Language Model Generations
- ResearchQA: Evaluating Scholarly Question Answering at Scale Across 75 Fields with Survey-Mined Questions and Rubrics
- UI-Bench: A Benchmark for Evaluating Design Capabilities of AI Text-to-App Tools
- IntentionReasoner: Facilitating Adaptive LLM Safeguards through Intent Reasoning and Selective Query Refinement
- ReSURE: Regularizing Supervision Unreliability for Multi-turn Dialogue Fine-tuning
- HEAL: A Hypothesis-Based Preference-Aware Analysis Framework
- Generative Interfaces for Language Models
- Rethinking Caching for LLM Serving Systems: Beyond Traditional Heuristics
- DEPTH: Hallucination-Free Relation Extraction via Dependency-Aware Sentence Simplification and Two-tiered Hierarchical Refinement
- The illusion of a perfect metric: Why evaluating AI's words is harder than it looks
- Prompt-Based One-Shot Exact Length-Controlled Generation with LLMs
- Bridging Human and LLM Judgments: Understanding and Narrowing the Gap
- Diversity First, Quality Later: A Two-Stage Assumption for Language Model Alignment
- Exploring Safety Alignment Evaluation of LLMs in Chinese Mental Health Dialogues via LLM-as-Judge
- Data Selection for LLM Alignment Using Fine-Grained Preferences
- Enhancing Small LLM Alignment through Margin-Based Objective Modifications under Resource Constraints
- Context Misleads LLMs: The Role of Context Filtering in Maintaining Safe Alignment of LLMs
- Synthesizing scientific literature with retrieval-augmented language models
- DP-LLM: Runtime Model Adaptation with Dynamic Layer-wise Precision Assignment
- Llama-3.1-FoundationAI-SecurityLLM-8B-Instruct Technical Report
- Cascaded Information Disclosure for Generalized Evaluation of Problem Solving Capabilities
- CoT-Self-Instruct: Building high-quality synthetic prompts for reasoning and non-reasoning tasks
- P3: Prompts Promote Prompting
- Libra: Assessing and Improving Reward Model by Learning to Think
- Checklists Are Better Than Reward Models For Aligning Language Models
- AlphaAlign: Incentivizing Safety Alignment with Extremely Simplified Reinforcement Learning
- Benchmarking LLM Privacy Recognition for Social Robot Decision Making
- Inverse Reinforcement Learning Meets Large Language Model Post-Training: Basics, Advances, and Opportunities
- Apple Intelligence Foundation Language Models: Tech Report 2025
- ConfPO: Exploiting Policy Model Confidence for Critical Token Selection in Preference Optimization
- Mitigating Reward Over-optimization in Direct Alignment Algorithms with Importance Sampling
- MultiVox: A Benchmark for Evaluating Voice Assistants for Multimodal Interactions
- Towards Concise and Adaptive Thinking in Large Reasoning Models: A Survey
- CompassJudger-2: Towards Generalist Judge Model via Verifiable Rewards
- Self-Improving Model Steering
- Quantile Reward Policy Optimization: Alignment with Pointwise Regression and Exact Partition Functions
- Stable Preference Optimization: A Bilevel Approach to Catastrophic Preference Shift
- The Delta Learning Hypothesis: Preference Tuning on Weak Data can Yield Strong Gains
- ContextASR-Bench: A Massive Contextual Speech Recognition Benchmark
- Cognitive World Model for Progressive BDI/E Trajectory Evaluation of Conversational Agents
- Hy-MultiTurn: A Six-Dimensional Benchmark for Deep Multi-Turn Dialogue Understanding
- Discrete Diffusion Trajectory Alignment via Stepwise Decomposition
- Nile-Chat: Egyptian Language Models for Arabic and Latin Scripts
- Meta SecAlign: A Secure Foundation LLM Against Prompt Injection Attacks
- On Reasoning Strength Planning in Large Reasoning Models
- FineInstructions: Scaling Synthetic Instructions to Pre-Training Scale
- Explicit Preference Optimization: No Need for an Implicit Reward Model
- Bridging Offline and Online Reinforcement Learning for LLMs
- MoE-GPS: Guidlines for Prediction Strategy for Dynamic Expert Duplication in MoE Load Balancing
- Shrinking the Generation-Verification Gap with Weak Verifiers
- Aligning Frozen LLMs by Reinforcement Learning: An Iterative Reweight-then-Optimize Approach
- DCRM: A Heuristic to Measure Response Pair Quality in Preference Optimization
- An Empirical Study of LLM-as-a-Judge: How Design Choices Impact Evaluation Reliability
- Verifying the Verifiers: Unveiling Pitfalls and Potentials in Fact Verifiers
- Flexible Realignment of Language Models
- Adaptive Batch-Wise Sample Scheduling for Direct Preference Optimization
- Implicit Reward as the Bridge: A Unified View of SFT and DPO Connections
- AI Flow: Perspectives, Scenarios, and Approaches
- AlphaSteer: Learning Refusal Steering with Principled Null-Space Constraint
- OneEval: Benchmarking LLM Knowledge-intensive Reasoning over Diverse Knowledge Bases
- Debiasing Online Preference Learning via Preference Feature Preservation
- dots.llm1 Technical Report
- Audio-Aware Large Language Models as Judges for Speaking Styles
- From Real to Synthetic: Synthesizing Millions of Diversified and Complicated User Instructions with Attributed Grounding
- Robust Preference Optimization via Dynamic Target Margins
- Aligning Large Language Models with Implicit Preferences from User-Generated Content
- T-SHIRT: Token-Selective Hierarchical Data Selection for Instruction Tuning
- Doubly Robust Alignment for Large Language Models
- AutoMixAlign: Adaptive Data Mixing for Multi-Task Preference Optimization in LLMs
- FACE: A Fine-grained Reference Free Evaluator for Conversational Recommender Systems
- Adversarial Preference Learning for Robust LLM Alignment
- Proximalized Preference Optimization for Diverse Feedback Types: A Decomposed Perspective on DPO
- Text2Grad: Reinforcement Learning from Natural Language Feedback
- A Course Correction in Steerability Evaluation: Revealing Miscalibration and Side Effects in LLMs
- Leveraging Importance Sampling to Detach Alignment Modules from Large Language Models
- Risk-aware Direct Preference Optimization under Nested Risk Measure
- Learning to Reason without External Rewards
- Improving Model Alignment Through Collective Intelligence of Open-Source LLMS
- Assistant-Guided Mitigation of Teacher Preference Bias in LLM-as-a-Judge
- The Price of Format: Diversity Collapse in LLMs
- Knowledge Grafting of Large Language Models
- OpenReview Should be Protected and Leveraged as a Community Asset for Research in the Era of Large Language Models
- TAG-INSTRUCT: Controlled Instruction Complexity Enhancement through Structure-based Augmentation
- Speechless: Speech Instruction Training Without Speech for Low Resource Languages
- PMPO: Probabilistic Metric Prompt Optimization for Small and Large Language Models
- Robustifying Vision-Language Models via Dynamic Token Reweighting
- Guiding Giants: Lightweight Controllers for Weighted Activation Steering in LLMs
- Latent Principle Discovery for Language Model Self-Improvement
- Be Careful When Fine-tuning On Open-Source LLMs: Your Fine-tuning Data Could Be Secretly Stolen!
- SLMEval: Entropy-Based Calibration for Human-Aligned Evaluation of Large Language Models
- TinyV: Reducing False Negatives in Verification Improves RL for LLM Reasoning
- InfiGFusion: Graph-on-Logits Distillation via Efficient Gromov-Wasserstein for Model Fusion
- Think Only When You Need with Large Hybrid-Reasoning Models
- Bias Fitting to Mitigate Length Bias of Reward Model in RLHF
- Decentralized Arena: Towards Democratic and Scalable Automatic Evaluation of Language Models
- Rethinking Predictive Modeling for LLM Routing: When Simple kNN Beats Complex Learned Routers
- R3: Robust Rubric-Agnostic Reward Models
- SGDPO: Self-Guided Direct Preference Optimization for Language Model Alignment
- ESC-Judge: A Framework for Comparing Emotional Support Conversational Agents
- AdAEM: An Adaptively and Automated Extensible Measurement of LLMs' Value Difference
- XtraGPT: Context-Aware and Controllable Academic Paper Revision via Human-AI Collaboration
- HelpSteer3-Preference: Open Human-Annotated Preference Data across Diverse Tasks and Languages
- WorldPM: Scaling Human Preference Modeling
- MorphMark: Flexible Adaptive Watermarking for Large Language Models
- Direct Density Ratio Optimization: A Statistically Consistent Approach to Aligning Large Language Models
- SCAN: Structured Capability Assessment and Navigation for LLMs
- Assessing Robustness to Spurious Correlations in Post-Training Language Models
- DialDefer: A Framework for Detecting and Mitigating LLM Dialogic Deference
- R-Bench: Graduate-level Multi-disciplinary Benchmarks for LLM & MLLM Complex Reasoning Evaluation
- Cer-Eval: Certifiable and Cost-Efficient Evaluation Framework for LLMs
- Quo Vadis, World Modeling?
- JudgeArena: A Unified Framework for Reproducible LLM-Judge Evaluation
- Invisible failures in human-AI interactions
- Evaluating Financial Intelligence in Large Language Models: Benchmarking SuperInvesting AI with LLM Engines
- Generalistic or Specific Embeddings, Which is Better? An Empirical Study on Search for Clinical Coding in Non-English Languages
- In LLM Reasoning, there is Irrationality on top of Value Misalignment
- MDIA: A Multi-Agent Diagnostic Intelligence Pipeline on HealthBench Professional
- Meeseeks: A Feedback-Driven, Iterative Self-Correction Benchmark evaluating LLMs' Instruction Following Capability
- General Preference Reinforcement Learning
- Automatic Legal Writing Evaluation of LLMs
- Robust AI Evaluation through Maximal Lotteries
- Anyprefer: An Agentic Framework for Preference Data Synthesis
- Toward Generalizable Evaluation in the LLM Era: A Survey Beyond Benchmarks
- AI Agents Need Memory Control Over More Context
- Do Chatbot LLMs Talk Too Much? The YapBench Benchmark
- Learning Explainable Dense Reward Shapes via Bayesian Optimization
- The Bitter Lesson Learned from 2,000+ Multilingual Benchmarks
- Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators
- DistilQwen2.5: Industrial Practices of Training Distilled Open Lightweight Language Models
- Pairwise or Pointwise? Evaluating Feedback Protocols for Bias in LLM-Based Evaluation
- ParaPO: Aligning Language Models to Reduce Verbatim Reproduction of Pre-training Data
- MIG: Automatic Data Selection for Instruction Tuning by Maximizing Information Gain in Semantic Space
- ImPart: Importance-Aware Delta-Sparsification for Improved Model Compression and Merging in LLMs
- Efficient MAP Estimation of LLM Judgment Performance with Prior Transfer
- Improving Instruct Models for Free: A Study on Partial Adaptation
- Learning from Reference Answers: Versatile Language Model Alignment without Binary Human Preference Data
- AdaSteer: Your Aligned LLM is Inherently an Adaptive Jailbreak Defender
- Leveraging Robust Optimization for LLM Alignment under Distribution Shifts
Related