vix.ing · top · new · best · stats · spec

Circuit-Aware Reward Training: A Mechanistic Framework for Longtail Robustness in RLHF

2025/09/29 by Liu, Jing
Engineering · #Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Safety Systems Engineering in Autonomy

paper · pdf · doi:10.48550/arxiv.2509.24713

openalex publication_date 2025/09/29 · openalex created_date 2025/10/19 · openalex updated_date 2026/07/28

Abstract

Reinforcement Learning from Human Feedback (RLHF) reward models exhibit systematic failures on longtail distributions, leading to reward hacking and misalignment. We propose a mechanistic interpretability framework that identifies specialized neural circuits responsible for rare-event processing in reward models. Drawing from recent advances showing distributed specialization for rare tokens in language models\citepliu2025no, liu2025emergent, we hypothesize that reward models also develop functionally distinct circuits for longtail scenarios. Our theoretical framework establishes formal connections between circuit specialization, reward generalization bounds, and longtail performance. We introduce Circuit-Aware Reward Training (CART), which uses circuit analysis to guide data augmentation, regularization, and ensemble strategies. This approach provides both theoretical insights into reward model failures and practical interventions for improving longtail robustness.

Citations

Related