vix.ing · top · new · best · stats

Learning Visually Guided Latent Actions for Assistive Teleoperation

2021/05/02 by Siddharth Karamcheti, Karamcheti, Siddharth, Albert J. Zhai +5 · 3 citations
Computer Science · Engineering · Neuroscience · #Artificial Intelligence (cs.AI) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #FOS: Electrical engineering #Human Pose and Action Recognition #Human-Computer Interaction (cs.HC) #Robot Manipulation and Learning #Robotics (cs.RO) #Systems and Control (eess.SY) #Tactile and Sensory Interactions #cs.AI #cs.CV #cs.HC #cs.RO #cs.SY #eess.SY #electronic engineering #information engineering

paper · pdf · doi:10.48550/arxiv.2105.00580

Accepted at Learning for Dynamics and Control (L4DC) 2021. 12 pages, 4 figures

arxiv created 2021/05/02 · openalex publication_date 2021/05/02 · arxiv updated 2021/05/04 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28

Abstract

It is challenging for humans -- particularly those living with physical disabilities -- to control high-dimensional, dexterous robots. Prior work explores learning embedding functions that map a human's low-dimensional inputs (e.g., via a joystick) to complex, high-dimensional robot actions for assistive teleoperation; however, a central problem is that there are many more high-dimensional actions than available low-dimensional inputs. To extract the correct action and maximally assist their human controller, robots must reason over their context: for example, pressing a joystick down when interacting with a coffee cup indicates a different action than when interacting with knife. In this work, we develop assistive robots that condition their latent embeddings on visual inputs. We explore a spectrum of visual encoders and show that incorporating object detectors pretrained on small amounts of cheap, easy-to-collect structured data enables i) accurately and robustly recognizing the current context and ii) generalizing control embeddings to new objects and tasks. In user studies with a high-dimensional physical robot arm, participants leverage this approach to perform new tasks with unseen objects. Our results indicate that structured visual representations improve few-shot performance and are subjectively preferred by users.

Citations

Cited by

Related