2025/07/11 by Karim Galliamov, Ivan Titov, Galliamov, Karim +3
Computer Science · #FOS: Computer and information sciences #Gaze Tracking and Assistive Technology #Machine Learning (cs.LG)
paper · pdf · doi:10.48550/arxiv.2507.09016
openalex publication_date 2025/07/11 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Reinforcement Learning from Human Feedback (RLHF) aligns language models with human preferences but is computationally expensive. We explore two approaches that leverage human gaze modeling to enhance RLHF: (1) gaze-aware reward models and (2) gaze-based distribution of sparse rewards at token level. Our experiments demonstate that gaze-informed RLHF achieves faster convergence while maintaining or slightly improving performance, thus, reducing computational costs during policy optimization. These results show that human gaze provides a valuable and underused signal for policy optimization, pointing to a promising direction for improving RLHF efficiency.