2021/03/01 by Vinay Kaushik, Kartik Jindgar, Kaushik, Vinay +3
Computer Science · Engineering · Mathematics · #Advanced Vision and Imaging #Artificial intelligence #Computer Vision and Pattern Recognition (cs.CV) #Computer science #Computer vision #Consistency (knowledge bases) #Deep learning #Depth map #Depth perception #FOS: Computer and information sciences #Generalization #Ground truth #Image (mathematics) #Image Processing Techniques and Applications #Machine learning #Mathematics #Monocular #Optical measurement and interference techniques #Pattern recognition (psychology) #Perception #cs.CV
paper · pdf · doi:10.48550/arxiv.2103.00853
published in arXiv (Cornell University) (Cornell University) · 8 pages
arxiv created 2021/03/01 · openalex publication_date 2021/03/01 · arxiv updated 2021/03/02 · openalex created_date 2022/07/25 · openalex updated_date 2026/08/06
Self-supervised learning of depth has been a highly studied topic of research\nas it alleviates the requirement of having ground truth annotations for\npredicting depth. Depth is learnt as an intermediate solution to the task of\nview synthesis, utilising warped photometric consistency. Although it gives\ngood results when trained using stereo data, the predicted depth is still\nsensitive to noise, illumination changes and specular reflections. Also,\nocclusion can be tackled better by learning depth from a single camera. We\npropose ADAA, utilising depth augmentation as depth supervision for learning\naccurate and robust depth. We propose a relational self-attention module that\nlearns rich contextual features and further enhances depth results. We also\noptimize the auto-masking strategy across all losses by enforcing L1\nregularisation over mask. Our novel progressive training strategy first learns\ndepth at a lower resolution and then progresses to the original resolution with\nslight training. We utilise a ResNet18 encoder, learning features for\nprediction of both depth and pose. We evaluate our predicted depth on the\nstandard KITTI driving dataset and achieve state-of-the-art results for\nmonocular depth estimation whilst having significantly lower number of\ntrainable parameters in our deep learning framework. We also evaluate our model\non Make3D dataset showing better generalization than other methods.\n