vix.ing · top · new · best · stats · spec

SooHwan Eom

  1. TLCR: Token-Level Continuous Reward for Fine-grained Reinforcement Learning from Human Feedback
    2024/07/23 by Eunseop Yoon, Hee Suk Yoon, Yoon, Eunseop +17 · 12 citations
    Computer Science · Engineering · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Muscle activation and electromyography studies #Reinforcement Learning in Robotics