2019/08/05 by Nusrah Hussain, Engin Erzin, Hussain, Nusrah +5
Computer Science · Psychology · #Reinforcement Learning in Robotics #Social Robot Interaction and HRI
paper · pdf · doi:10.48550/arxiv.1908.01618
We present a novel method for training a social robot to generate\nbackchannels during human-robot interaction. We address the problem within an\noff-policy reinforcement learning framework, and show how a robot may learn to\nproduce non-verbal backchannels like laughs, when trained to maximize the\nengagement and attention of the user. A major contribution of this work is the\nformulation of the problem as a Markov decision process (MDP) with states\ndefined by the speech activity of the user and rewards generated by quantified\nengagement levels. The problem that we address falls into the class of\napplications where unlimited interaction with the environment is not possible\n(our environment being a human) because it may be time-consuming, costly,\nimpracticable or even dangerous in case a bad policy is executed. Therefore, we\nintroduce deep Q-network (DQN) in a batch reinforcement learning framework,\nwhere an optimal policy is learned from a batch data collected using a more\ncontrolled policy. We suggest the use of human-to-human dyadic interaction\ndatasets as a batch of trajectories to train an agent for engaging\ninteractions. Our experiments demonstrate the potential of our method to train\na robot for engaging behaviors in an offline manner.\n