2019/10/10 by Erdem Bıyık, Bıyık, Erdem, Malayandi Palan +7 · 7 citations
Computer Science · Decision Sciences · #Advanced Bandit Algorithms Research #Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning and Algorithms #Reinforcement Learning in Robotics #Robotics (cs.RO)
paper · pdf · doi:10.48550/arxiv.1910.04365
openalex publication_date 2019/10/10 · openalex created_date 2022/07/28 · openalex updated_date 2026/08/04
Robots can learn the right reward function by querying a human expert.\nExisting approaches attempt to choose questions where the robot is most\nuncertain about the human's response; however, they do not consider how easy it\nwill be for the human to answer! In this paper we explore an information gain\nformulation for optimally selecting questions that naturally account for the\nhuman's ability to answer. Our approach identifies questions that optimize the\ntrade-off between robot and human uncertainty, and determines when these\nquestions become redundant or costly. Simulations and a user study show our\nmethod not only produces easy questions, but also ultimately results in faster\nreward learning.\n