2019/07/27 by Michelle A. Lee, Lee, Michelle A., Yuke Zhu +15 · 27 citations
Computer Science · Engineering · Neuroscience · #Reinforcement Learning in Robotics #Robot Manipulation and Learning #Tactile and Sensory Interactions #cs.LG #cs.RO
paper · pdf · doi:10.48550/arxiv.1907.13098
arXiv admin note: substantial text overlap with arXiv:1810.10191
arxiv created 2019/07/28 · arxiv updated 2019/07/31
Contact-rich manipulation tasks in unstructured environments often require both haptic and visual feedback. It is non-trivial to manually design a robot controller that combines these modalities which have very different characteristics. While deep reinforcement learning has shown success in learning control policies for high-dimensional inputs, these algorithms are generally intractable to deploy on real robots due to sample complexity. In this work, we use self-supervision to learn a compact and multimodal representation of our sensory inputs, which can then be used to improve the sample efficiency of our policy learning. Evaluating our method on a peg insertion task, we show that it generalizes over varying geometries, configurations, and clearances, while being robust to external perturbations. We also systematically study different self-supervised learning objectives and representation learning architectures. Results are presented in simulation and on a physical robot.