2019/04/17 by Robin Kreuzig, Matthias Ochs, Kreuzig, Robin +3 · 1 citation
Computer Science · #Human Pose and Action Recognition #Advanced Vision and Imaging #Video Surveillance and Tracking Methods
paper · pdf · doi:10.48550/arxiv.1904.08105
Classical monocular vSLAM/VO methods suffer from the scale ambiguity problem.\nHybrid approaches solve this problem by adding deep learning methods, for\nexample by using depth maps which are predicted by a CNN. We suggest that it is\nbetter to base scale estimation on estimating the traveled distance for a set\nof subsequent images. In this paper, we propose a novel end-to-end many-to-one\ntraveled distance estimator. By using a deep recurrent convolutional neural\nnetwork (RCNN), the traveled distance between the first and last image of a set\nof consecutive frames is estimated by our DistanceNet. Geometric features are\nlearned in the CNN part of our model, which are subsequently used by the RNN to\nlearn dynamics and temporal information. Moreover, we exploit the natural order\nof distances by using ordinal regression to predict the distance. The\nevaluation on the KITTI dataset shows that our approach outperforms current\nstate-of-the-art deep learning pose estimators and classical mono vSLAM/VO\nmethods in terms of distance prediction. Thus, our DistanceNet can be used as a\ncomponent to solve the scale problem and help improve current and future\nclassical mono vSLAM/VO methods.\n