2019/12/11 by Glen Berseth, Berseth, Glen, Daniel Geng +11 · 1 citation
Computer Science · Neuroscience · Social Sciences · #Artificial Intelligence (cs.AI) #Evolutionary Game Theory and Cooperation #FOS: Computer and information sciences #G.3 #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Neural dynamics and brain function #Reinforcement Learning in Robotics
paper · pdf · doi:10.48550/arxiv.1912.05510
openalex publication_date 2019/12/11 · openalex created_date 2022/07/26 · openalex updated_date 2026/07/28
Every living organism struggles against disruptive environmental forces to\ncarve out and maintain an orderly niche. We propose that such a struggle to\nachieve and preserve order might offer a principle for the emergence of useful\nbehaviors in artificial agents. We formalize this idea into an unsupervised\nreinforcement learning method called surprise minimizing reinforcement learning\n(SMiRL). SMiRL alternates between learning a density model to evaluate the\nsurprise of a stimulus, and improving the policy to seek more predictable\nstimuli. The policy seeks out stable and repeatable situations that counteract\nthe environment's prevailing sources of entropy. This might include avoiding\nother hostile agents, or finding a stable, balanced pose for a bipedal robot in\nthe face of disturbance forces. We demonstrate that our surprise minimizing\nagents can successfully play Tetris, Doom, control a humanoid to avoid falls,\nand navigate to escape enemies in a maze without any task-specific reward\nsupervision. We further show that SMiRL can be used together with standard task\nrewards to accelerate reward-driven learning.\n