2021/11/29 by Assem Sadek, Sadek, Assem, Guillaume Bono +5 · 1 citation
Computer Science · Engineering · #Advanced Image and Video Retrieval Techniques #Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Multimodal Machine Learning Applications #Robotics (cs.RO) #Robotics and Sensor-Based Localization
paper · pdf · doi:10.48550/arxiv.2111.14666
openalex publication_date 2021/11/29 · openalex created_date 2022/11/27 · openalex updated_date 2026/07/28
Visual navigation by mobile robots is classically tackled through SLAM plus\noptimal planning, and more recently through end-to-end training of policies\nimplemented as deep networks. While the former are often limited to waypoint\nplanning, but have proven their efficiency even on real physical environments,\nthe latter solutions are most frequently employed in simulation, but have been\nshown to be able learn more complex visual reasoning, involving complex\nsemantical regularities. Navigation by real robots in physical environments is\nstill an open problem. End-to-end training approaches have been thoroughly\ntested in simulation only, with experiments involving real robots being\nrestricted to rare performance evaluations in simplified laboratory conditions.\nIn this work we present an in-depth study of the performance and reasoning\ncapacities of real physical agents, trained in simulation and deployed to two\ndifferent physical environments. Beyond benchmarking, we provide insights into\nthe generalization capabilities of different agents training in different\nconditions. We visualize sensor usage and the importance of the different types\nof signals. We show, that for the PointGoal task, an agent pre-trained on wide\nvariety of tasks and fine-tuned on a simulated version of the target\nenvironment can reach competitive performance without modelling any sim2real\ntransfer, i.e. by deploying the trained agent directly from simulation to a\nreal physical robot.\n