2018/09/21 by Luca Caltagirone, Mauro Bellone, Caltagirone, Luca +5 · 4 citations
Computer Science · Engineering · Environmental Science · #Advanced Neural Network Applications #Autonomous Vehicle Technology and Safety #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Infrastructure Maintenance and Monitoring #Remote Sensing and LiDAR Applications
paper · pdf · doi:10.48550/arxiv.1809.07941
openalex publication_date 2018/09/21 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
In this work, a deep learning approach has been developed to carry out road\ndetection by fusing LIDAR point clouds and camera images. An unstructured and\nsparse point cloud is first projected onto the camera image plane and then\nupsampled to obtain a set of dense 2D images encoding spatial information.\nSeveral fully convolutional neural networks (FCNs) are then trained to carry\nout road detection, either by using data from a single sensor, or by using\nthree fusion strategies: early, late, and the newly proposed cross fusion.\nWhereas in the former two fusion approaches, the integration of multimodal\ninformation is carried out at a predefined depth level, the cross fusion FCN is\ndesigned to directly learn from data where to integrate information; this is\naccomplished by using trainable cross connections between the LIDAR and the\ncamera processing branches.\n To further highlight the benefits of using a multimodal system for road\ndetection, a data set consisting of visually challenging scenes was extracted\nfrom driving sequences of the KITTI raw data set. It was then demonstrated\nthat, as expected, a purely camera-based FCN severely underperforms on this\ndata set. A multimodal system, on the other hand, is still able to provide high\naccuracy. Finally, the proposed cross fusion FCN was evaluated on the KITTI\nroad benchmark where it achieved excellent performance, with a MaxF score of\n96.03%, ranking it among the top-performing approaches.\n