vix.ing · top · new · best · stats · spec

Learning Generalizable Robotic Reward Functions from "In-The-Wild" Human\n Videos

2021/03/31 by Annie S. Chen, Suraj Nair, Chen, Annie S. +3 · 14 citations
Computer Science · #Artificial Intelligence (cs.AI) #Computer Vision and Pattern Recognition (cs.CV) #Domain Adaptation and Few-Shot Learning #FOS: Computer and information sciences #Machine Learning (cs.LG) #Multimodal Machine Learning Applications #Reinforcement Learning in Robotics #Robotics (cs.RO)

paper · pdf · doi:10.48550/arxiv.2103.16817

openalex publication_date 2021/03/31 · openalex created_date 2022/07/25 · openalex updated_date 2026/07/28

Abstract

We are motivated by the goal of generalist robots that can complete a wide\nrange of tasks across many environments. Critical to this is the robot's\nability to acquire some metric of task success or reward, which is necessary\nfor reinforcement learning, planning, or knowing when to ask for help. For a\ngeneral-purpose robot operating in the real world, this reward function must\nalso be able to generalize broadly across environments, tasks, and objects,\nwhile depending only on on-board sensor observations (e.g. RGB images). While\ndeep learning on large and diverse datasets has shown promise as a path towards\nsuch generalization in computer vision and natural language, collecting high\nquality datasets of robotic interaction at scale remains an open challenge. In\ncontrast, "in-the-wild" videos of humans (e.g. YouTube) contain an extensive\ncollection of people doing interesting tasks across a diverse range of\nsettings. In this work, we propose a simple approach, Domain-agnostic Video\nDiscriminator (DVD), that learns multitask reward functions by training a\ndiscriminator to classify whether two videos are performing the same task, and\ncan generalize by virtue of learning from a small amount of robot data with a\nbroad dataset of human videos. We find that by leveraging diverse human\ndatasets, this reward function (a) can generalize zero shot to unseen\nenvironments, (b) generalize zero shot to unseen tasks, and (c) can be combined\nwith visual model predictive control to solve robotic manipulation tasks on a\nreal WidowX200 robot in an unseen environment from a single human demo.\n

Cited by

Related