2020/07/29 by Cédric Picron, Punarjay Chakravarty, Picron, Cédric +5
Computer Science · Engineering · #Advanced Neural Network Applications #Advanced Vision and Imaging #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Robotics and Sensor-Based Localization
paper · pdf · doi:10.48550/arxiv.2007.14812
openalex publication_date 2020/07/29 · openalex created_date 2022/10/01 · openalex updated_date 2026/07/28
The estimation of the orientation of an observed vehicle relative to an\nAutonomous Vehicle (AV) from monocular camera data is an important building\nblock in estimating its 6 DoF pose. Current Deep Learning based solutions for\nplacing a 3D bounding box around this observed vehicle are data hungry and do\nnot generalize well. In this paper, we demonstrate the use of monocular visual\nodometry for the self-supervised fine-tuning of a model for orientation\nestimation pre-trained on a reference domain. Specifically, while transitioning\nfrom a virtual dataset (vKITTI) to nuScenes, we recover up to 70% of the\nperformance of a fully supervised method. We subsequently demonstrate an\noptimization-based monocular 3D bounding box detector built on top of the\nself-supervised vehicle orientation estimator without the requirement of\nexpensive labeled data. This allows 3D vehicle detection algorithms to be\nself-trained from large amounts of monocular camera data from existing\ncommercial vehicle fleets.\n