2025/02/15 by Hongyu Yang, Yang, Hongyu, Qi Zhao +5
Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Data Management and Algorithms #FOS: Computer and information sciences #Rough Sets and Fuzzy Logic
paper · pdf · doi:10.48550/arxiv.2502.12189
openalex publication_date 2025/02/15 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Reinforcement Learning from Human Feedback and its variants excel in aligning with human intentions to generate helpful, harmless, and honest responses. However, most of them rely on costly human-annotated pairwise comparisons for supervised alignment, which is not suitable for list-level scenarios, such as community question answering. Additionally, human preferences are influenced by multiple intrinsic factors in responses, leading to decision-making inconsistencies. Therefore, we propose Self-supervised Attribute-aware dynamic preference ranking, called \shortname. It quantifies preference differences between responses based on Attribute-Perceptual Distance Factors (APDF) and dynamically determines the list-wise alignment order. Furthermore, it achieves fine-grained preference difference learning and enables precise alignment with the optimal one. We specifically constructed a challenging code preference dataset named StaCoCoQA, and introduced more cost-effective and scalable preference evaluation metrics: PrefHit and PrefRecall. Extensive experimental results show that SeAdpra exhibits superior performance and generalizability on both StaCoCoQA and preference datasets from eight popular domains.