2017/05/21 by Min Xu, Xu, Min
Computer Science · Decision Sciences · Neuroscience · #Advanced Bandit Algorithms Research #Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Neural dynamics and brain function #Reinforcement Learning in Robotics #cs.AI
paper · pdf · doi:10.48550/arxiv.1705.07460
4 pages, 1 figure
arxiv created 2017/05/21 · openalex publication_date 2017/05/21 · arxiv updated 2017/05/23 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
For most reinforcement learning approaches, the learning is performed by maximizing an accumulative reward that is expectedly and manually defined for specific tasks. However, in real world, rewards are emergent phenomena from the complex interactions between agents and environments. In this paper, we propose an implicit generic reward model for reinforcement learning. Unlike those rewards that are manually defined for specific tasks, such implicit reward is task independent. It only comes from the deviation from the agents' previous experiences.