2019/09/19 by Tianwei Shen, Lei Zhou, Shen, Tianwei +13 · 1 citation
Computer Science · Engineering · #Advanced Vision and Imaging #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Optical measurement and interference techniques #Robotics and Sensor-Based Localization
paper · pdf · doi:10.48550/arxiv.1909.09115
openalex publication_date 2019/09/19 · openalex created_date 2019/09/26 · openalex updated_date 2026/07/28
The self-supervised learning of depth and pose from monocular sequences provides an attractive solution by using the photometric consistency of nearby frames as it depends much less on the ground-truth data. In this paper, we address the issue when previous assumptions of the self-supervised approaches are violated due to the dynamic nature of real-world scenes. Different from handling the noise as uncertainty, our key idea is to incorporate more robust geometric quantities and enforce internal consistency in the temporal image sequence. As demonstrated on commonly used benchmark datasets, the proposed method substantially improves the state-of-the-art methods on both depth and relative pose estimation for monocular image sequences, without adding inference overhead.