RCTDistill: Cross-Modal Knowledge Distillation Framework for Radar-Camera 3D Object Detection with Temporal Fusion
2025/09/22 by Bang, Geonho, Seong, Minjae, Kim, Jisong +5
#Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences
paper · doi:10.48550/arxiv.2509.17712
Abstract
Radar-camera fusion methods have emerged as a cost-effective approach for 3D object detection but still lag behind LiDAR-based methods in performance. Recent works have focused on employing temporal fusion and Knowledge Distillation (KD) strategies to overcome these limitations. However, existing approaches have not sufficiently accounted for uncertainties arising from object motion or sensor-specific errors inherent in radar and camera modalities. In this work, we propose RCTDistill, a novel cross-modal KD method based on temporal fusion, comprising three key modules: Range-Azimuth Knowledge Distillation (RAKD), Temporal Knowledge Distillation (TKD), and Region-Decoupled Knowledge Distillation (RDKD). RAKD is designed to consider the inherent errors in the range and azimuth directions, enabling effective knowledge transfer from LiDAR features to refine inaccurate BEV representations. TKD mitigates temporal misalignment caused by dynamic objects by aligning historical radar-camera BEV features with current LiDAR representations. RDKD enhances feature discrimination by distilling relational knowledge from the teacher model, allowing the student to differentiate foreground and background features. RCTDistill achieves state-of-the-art radar-camera fusion performance on both the nuScenes and View-of-Delft (VoD) datasets, with the fastest inference speed of 26.2 FPS.
Citations
- Toward Real-world BEV Perception: Depth Uncertainty Estimation via Gaussian Splatting
- RCTrans: Radar-Camera Transformer via Radar Densifier and Sequential Decoder for 3D Object Detection
- HGSFusion: Radar-Camera Fusion with Hybrid Generation and Synchronization for 3D Object Detection
- SpaRC: Sparse Radar-Camera Fusion for 3D Object Detection
- CRT-Fusion: Camera, Radar, Temporal Fusion Using Motion Information for 3D Object Detection
- LEROjD: Lidar Extended Radar-Only Object Detection
- CRKD: Enhanced Camera-Radar Object Detection with Cross-modality Knowledge Distillation
- RCBEVDet: Radar-camera Fusion in Bird's Eye View for 3D Object Detection
- RadarDistill: Boosting Radar-based Object Detection Performance via Knowledge Distillation from LiDAR Features
- UniRepLKNet: A Universal Perception Large-Kernel ConvNet for Audio, Video, Point Cloud, Time-Series and Image Recognition
- Leveraging Vision-Centric Multi-Modal Expertise for 3D Object Detection
- DistillBEV: Boosting Multi-Camera 3D Object Detection with Cross-Modal Knowledge Distillation
- RCM-Fusion: Radar-Camera Multi-Level Fusion for 3D Object Detection
- CRN: Camera Radar Net for Accurate, Robust, Efficient 3D Perception
- SimDistill: Simulated Multi-modal Distillation for BEV 3D Object Detection
- UniDistill: A Universal Cross-Modality Knowledge Distillation Framework for 3D Object Detection in Bird's-Eye View
- X3KD: Knowledge Distillation Across Modalities, Tasks and Stages for Multi-Camera 3D Object Detection
- BEVDistill: Cross-Modal BEV Distillation for Multi-View 3D Object Detection
- LidarAugment: Searching for Scalable 3D LiDAR Data Augmentations
- Time Will Tell: New Outlooks and A Baseline for Temporal Multi-View 3D Object Detection
- Bridging the View Disparity Between Radar and Camera Features for Multi-modal Fusion 3D Object Detection
- BEVDepth: Acquisition of Reliable Depth for Multi-view 3D Object Detection
- Unifying Voxel-based Representation with Transformer for 3D Object Detection
- Boosting 3D Object Detection by Simulating Multimodality on Point Clouds
- BEVFusion: Multi-Task Multi-Sensor Fusion with Unified Bird's-Eye View Representation
- BEVFormer: Learning Bird's-Eye-View Representation from Multi-Camera Images via Spatiotemporal Transformers
- PETR: Position Embedding Transformation for Multi-View 3D Object\n Detection
- MonoDistill: Learning Spatial Features for Monocular 3D Object Detection
- A ConvNet for the 2020s
- A ConvNet for the 2020s
- BEVDet: High-performance Multi-camera 3D Object Detection in Bird-Eye-View
- DETR3D: 3D Object Detection from Multi-view Images via 3D-to-2D Queries
- LIGA-Stereo: Learning LiDAR Geometry Aware Representations for Stereo-based 3D Detector
- Lift, Splat, Shoot: Encoding Images From Arbitrary Camera Rigs by\n Implicitly Unprojecting to 3D
- Center-based 3D Object Detection and Tracking
- Class-balanced Grouping and Sampling for Point Cloud 3D Object Detection
- nuScenes: A multimodal dataset for autonomous driving
- PointPillars: Fast Encoders for Object Detection from Point Clouds
- Deep Ordinal Regression Network for Monocular Depth Estimation
- Deep Residual Learning for Image Recognition
Related