2020/09/22 by Ozan Unal, Luc Van Gool, Unal, Ozan +3 · 2 citations
Computer Science · Engineering · Environmental Science · #3D Shape Modeling and Analysis #Artificial intelligence #Class (philosophy) #Computer Vision and Pattern Recognition (cs.CV) #Computer science #Computer vision #Convolutional neural network #Deep learning #FOS: Computer and information sciences #Feature (linguistics) #Feature learning #Focus (optics) #Object (grammar) #Object detection #Orientation (vector space) #Path (computing) #Pipeline (software) #Point cloud #Pose #Remote Sensing and LiDAR Applications #Representation (politics) #Robotics and Sensor-Based Localization #Segmentation #Task (project management) #cs.CV
paper · pdf · doi:10.48550/arxiv.2009.10569
published in arXiv (Cornell University) (Cornell University) · Accepted at IEEE Winter Conference on Applications of Computer Vision 2021 (WACV'21)
openalex publication_date 2020/09/22 · arxiv created 2020/11/07 · arxiv updated 2020/11/10 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Point cloud semantic segmentation plays an essential role in autonomous driving, providing vital information about drivable surfaces and nearby objects that can aid higher level tasks such as path planning and collision avoidance. While current 3D semantic segmentation networks focus on convolutional architectures that perform great for well represented classes, they show a significant drop in performance for underrepresented classes that share similar geometric features. We propose a novel Detection Aware 3D Semantic Segmentation (DASS) framework that explicitly leverages localization features from an auxiliary 3D object detection task. By utilizing multitask training, the shared feature representation of the network is guided to be aware of per class detection features that aid tackling the differentiation of geometrically similar classes. We additionally provide a pipeline that uses DASS to generate high recall proposals for existing 2-stage detectors and demonstrate that the added supervisory signal can be used to improve 3D orientation estimation capabilities. Extensive experiments on both the SemanticKITTI and KITTI object datasets show that DASS can improve 3D semantic segmentation results of geometrically similar classes up to 37.8% IoU in image FOV while maintaining high precision bird's-eye view (BEV) detection results.