2025/01/13 by Yaqing Ding, Viktor Kocur, Ding, Yaqing +12 · 1 voice · 1 citation
Computer Science · Engineering · #Advanced Vision and Imaging #Optical measurement and interference techniques #Robotics and Sensor-Based Localization #cs.CV
paper · pdf · doi:10.48550/arxiv.2501.07742
openalex publication_date 2025/01/13 · openalex created_date 2025/01/17 · openalex updated_date 2026/07/28
Recent advances in monocular depth estimation methods (MDE) and their improved accuracy open new possibilities for their applications. In this paper, we investigate how monocular depth estimates can be used for relative pose estimation. In particular, we are interested in answering the question whether using MDEs improves results over traditional point-based methods. We propose a novel framework for estimating the relative pose of two cameras from point correspondences with associated monocular depths. Since depth predictions are typically defined up to an unknown scale or even both unknown scale and shift parameters, our solvers jointly estimate the scale or both the scale and shift parameters along with the relative pose. We derive efficient solvers considering different types of depths for three camera configurations: (1) two calibrated cameras, (2) two cameras with an unknown shared focal length, and (3) two cameras with unknown different focal lengths. Our new solvers outperform state-of-the-art depth-aware solvers in terms of speed and accuracy. In extensive real experiments on multiple datasets and with various MDEs, we discuss which depth-aware solvers are preferable in which situation. The code will be made publicly available.