2024/07/22 by Kaiwen Wang, Wang, Kaiwen, Rahul Kidambi +36 · 8 citations
Arts and Humanities · Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Linguistic research and analysis #Machine Learning (cs.LG) #Natural Language Processing Techniques #Syntax, Semantics, Linguistic Variation
paper · pdf · doi:10.48550/arxiv.2407.15762
openalex publication_date 2024/07/22 · openalex created_date 2024/09/26 · openalex updated_date 2026/07/28
Reward-based finetuning is crucial for aligning language policies with intended behaviors (e.g., creativity and safety). A key challenge is to develop steerable language models that trade-off multiple (conflicting) objectives in a flexible and efficient manner. This paper presents Conditional Language Policy (CLP), a general framework for finetuning language models on multiple objectives. Building on techniques from multi-task training and parameter-efficient finetuning, CLP learn steerable models that effectively trade-off conflicting objectives at inference time. Notably, this does not require training or maintaining multiple models to achieve different trade-offs between the objectives. Through extensive experiments and ablations on two summarization datasets, we show that CLP learns steerable language models that outperform and Pareto-dominate the existing approaches for multi-objective finetuning.