2022/06/30 by Wu Zheng, Zheng Wu, Mingxuan Hong +8 · 4 citations
Computer Science · #Advanced Neural Network Applications #Artificial intelligence #Boosting (machine learning) #Computer science #Computer vision #Consistency (knowledge bases) #Detector #Domain Adaptation and Few-Shot Learning #Feature (linguistics) #Focus (optics) #Geography #Human Pose and Action Recognition #Inference #Lidar #Modality (human–computer interaction) #Object detection #Optics #Pattern recognition (psychology) #Physics #Remote sensing #Voxel #cs.CV
paper · pdf · doi:10.48550/arxiv.2206.14971
published in arXiv (Cornell University) (Cornell University) · Published in CVPR 2022 as Oral
arxiv created 2022/06/30 · openalex publication_date 2022/06/30 · arxiv updated 2022/07/01 · openalex created_date 2025/10/10 · openalex updated_date 2026/08/04
This paper presents a new approach to boost a single-modality (LiDAR) 3D object detector by teaching it to simulate features and responses that follow a multi-modality (LiDAR-image) detector. The approach needs LiDAR-image data only when training the single-modality detector, and once well-trained, it only needs LiDAR data at inference. We design a novel framework to realize the approach: response distillation to focus on the crucial response samples and avoid the background samples; sparse-voxel distillation to learn voxel semantics and relations from the estimated crucial voxels; a fine-grained voxel-to-point distillation to better attend to features of small and distant objects; and instance distillation to further enhance the deep-feature consistency. Experimental results on the nuScenes dataset show that our approach outperforms all SOTA LiDAR-only 3D detectors and even surpasses the baseline LiDAR-image detector on the key NDS metric, filling 72% mAP gap between the single- and multi-modality detectors.