vix.ing · top · new · best · stats · spec

Assistant-Guided Mitigation of Teacher Preference Bias in LLM-as-a-Judge

2025/05/25 by Zhuo Liu, Moxin Li, Liu, Zhuo +7 · 2 citations
Business, Management and Accounting · Social Sciences · #Artificial Intelligence in Law #Computation and Language (cs.CL) #Dispute Resolution and Class Actions #FOS: Computer and information sciences #Legal Education and Practice Innovations

paper · pdf · doi:10.48550/arxiv.2505.19176

openalex publication_date 2025/05/25 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28

Abstract

LLM-as-a-Judge employs large language models (LLMs), such as GPT-4, to evaluate the quality of LLM-generated responses, gaining popularity for its cost-effectiveness and strong alignment with human evaluations. However, training proxy judge models using evaluation data generated by powerful teacher models introduces a critical yet previously overlooked issue: teacher preference bias, where the proxy judge model learns a biased preference for responses from the teacher model. To tackle this problem, we propose a novel setting that incorporates an additional assistant model, which is not biased toward the teacher model's responses, to complement the training data. Building on this setup, we introduce AGDe-Judge, a three-stage framework designed to debias from both the labels and feedbacks in the training data. Extensive experiments demonstrate that AGDe-Judge effectively reduces teacher preference bias while maintaining strong performance across six evaluation benchmarks. Code is available at https://github.com/Liuz233/AGDe-Judge.

Citations

Cited by

Related