vix.ing · top · new · best · stats

Rethinking the Evaluation of Alignment Methods: Insights into Diversity, Generalisation, and Safety

2025/09/16 by Denis Janiak, Janiak, Denis, Julia Moska +11
Decision Sciences · Pharmacology, Toxicology and Pharmaceutics · #Computation and Language (cs.CL) #Evaluation and Performance Assessment #FOS: Computer and information sciences #Machine Learning (cs.LG) #Pharmacy and Medical Practices

paper · pdf · doi:10.48550/arxiv.2509.12936

openalex publication_date 2025/09/16 · openalex created_date 2025/10/18 · openalex updated_date 2026/07/28

Abstract

Large language models (LLMs) require careful alignment to balance competing objectives - factuality, safety, conciseness, proactivity, and diversity. Existing studies focus on individual techniques or specific dimensions, lacking a holistic assessment of the inherent trade-offs. We propose a unified evaluation framework that compares LLM alignment methods (PPO, DPO, ORPO, KTO) across these five axes, using both in-distribution and out-of-distribution datasets. Leveraging a specialized LLM-as-Judge prompt, validated through human studies, we reveal that DPO and KTO excel in factual accuracy, PPO and DPO lead in safety, and PPO best balances conciseness with proactivity. Our findings provide insights into trade-offs of common alignment methods, guiding the development of more balanced and reliable LLMs.

Citations

Related