2019/10/21 by Xingyu Liu, Liu, Xingyu, Mengyuan Yan +3 · 16 citations
Computer Science · Engineering · Mathematics · #3D Shape Modeling and Analysis #Advanced Vision and Imaging #Artificial intelligence #Benchmark (surveying) #Computer Vision and Pattern Recognition (cs.CV) #Computer science #Construct (python library) #Data mining #Deep learning #FOS: Computer and information sciences #Geography #Grid #Human Pose and Action Recognition #Machine Learning (cs.LG) #Machine learning #Mathematics #Point (geometry) #Point cloud #Process (computing) #Representation (politics) #Robotics (cs.RO) #Segmentation #Sequence (biology) #Variety (cybernetics) #cs.CV #cs.LG #cs.RO
paper · pdf · doi:10.48550/arxiv.1910.09165
published in arXiv (Cornell University) (Cornell University) · ICCV 2019 (Oral)
openalex publication_date 2019/10/21 · arxiv created 2020/02/08 · arxiv updated 2020/02/11 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Understanding dynamic 3D environment is crucial for robotic agents and many other applications. We propose a novel neural network architecture called MeteorNet for learning representations for dynamic 3D point cloud sequences. Different from previous work that adopts a grid-based representation and applies 3D or 4D convolutions, our network directly processes point clouds. We propose two ways to construct spatiotemporal neighborhoods for each point in the point cloud sequence. Information from these neighborhoods is aggregated to learn features per point. We benchmark our network on a variety of 3D recognition tasks including action recognition, semantic segmentation and scene flow estimation. MeteorNet shows stronger performance than previous grid-based methods while achieving state-of-the-art performance on Synthia. MeteorNet also outperforms previous baseline methods that are able to process at most two consecutive point clouds. To the best of our knowledge, this is the first work on deep learning for dynamic raw point cloud sequences.