2020/04/10 by Lina Mezghani, Sainbayar Sukhbaatar, Mezghani, Lina +7 · 1 citation
Computer Science · #Advanced Image and Video Retrieval Techniques #Artificial Intelligence (cs.AI) #Computer Vision and Pattern Recognition (cs.CV) #Domain Adaptation and Few-Shot Learning #FOS: Computer and information sciences #Machine Learning (cs.LG) #Multimodal Machine Learning Applications #Robotics (cs.RO)
paper · pdf · doi:10.48550/arxiv.2004.04954
openalex publication_date 2020/04/10 · openalex created_date 2023/02/14 · openalex updated_date 2026/07/28
Learning to navigate in a realistic setting where an agent must rely solely\non visual inputs is a challenging task, in part because the lack of position\ninformation makes it difficult to provide supervision during training. In this\npaper, we introduce a novel approach for learning to navigate from image inputs\nwithout external supervision or reward. Our approach consists of three stages:\nlearning a good representation of first-person views, then learning to explore\nusing memory, and finally learning to navigate by setting its own goals. The\nmodel is trained with intrinsic rewards only so that it can be applied to any\nenvironment with image observations. We show the benefits of our approach by\ntraining an agent to navigate challenging photo-realistic environments from the\nGibson dataset with RGB inputs only.\n