2018/05/29 by Yusuf Aytar, Aytar, Yusuf, Tobias Pfaff +9 · 1 voice · 6 citations
Computer Science · Mathematics · Psychology · #Action (physics) #Artificial intelligence #Computer science #Construct (python library) #Human–computer interaction #Imitation #Multimodal Machine Learning Applications #Programming language #Psychology #Reinforcement Learning in Robotics #Reinforcement learning #Representation (politics) #cs.AI #cs.CV #cs.LG #stat.ML
paper · pdf · doi:10.48550/arxiv.1805.11592
published in arXiv (Cornell University) 31, 2930-2941 (Cornell University)
openalex publication_date 2018/05/29 · arxiv created 2018/11/30 · arxiv updated 2018/12/03 · openalex created_date 2025/10/10 · openalex updated_date 2026/08/05
Deep reinforcement learning methods traditionally struggle with tasks where environment rewards are particularly sparse. One successful method of guiding exploration in these domains is to imitate trajectories provided by a human demonstrator. However, these demonstrations are typically collected under artificial conditions, i.e. with access to the agent's exact environment setup and the demonstrator's action and reward trajectories. Here we propose a two-stage method that overcomes these limitations by relying on noisy, unaligned footage without access to such data. First, we learn to map unaligned videos from multiple sources to a common representation using self-supervised objectives constructed over both time and modality (i.e. vision and sound). Second, we embed a single YouTube video in this representation to construct a reward function that encourages an agent to imitate human gameplay. This method of one-shot imitation allows our agent to convincingly exceed human-level performance on the infamously hard exploration games Montezuma's Revenge, Pitfall! and Private Eye for the first time, even if the agent is not presented with any environment rewards.