2020/07/07 by Dongsu Zhang, Junha Chun, Zhang, Dongsu +7 · 13 citations
Computer Science · Engineering · #3D Shape Modeling and Analysis #Advanced Vision and Imaging #Artificial intelligence #Benchmark (surveying) #Cartography #Computer Vision and Pattern Recognition (cs.CV) #Computer science #Computer vision #Deep learning #Embedding #FOS: Computer and information sciences #Geography #Image (mathematics) #Machine learning #Metric (unit) #Noise (video) #Pattern recognition (psychology) #Robotics and Sensor-Based Localization #Scalability #Segmentation #cs.CV
paper · pdf · doi:10.48550/arxiv.2007.03169
published in arXiv (Cornell University) (Cornell University)
arxiv created 2020/07/07 · openalex publication_date 2020/07/07 · arxiv updated 2020/07/08 · openalex created_date 2020/07/10 · openalex updated_date 2026/08/05
We propose spatial semantic embedding network (SSEN), a simple, yet efficient algorithm for 3D instance segmentation using deep metric learning. The raw 3D reconstruction of an indoor environment suffers from occlusions, noise, and is produced without any meaningful distinction between individual entities. For high-level intelligent tasks from a large scale scene, 3D instance segmentation recognizes individual instances of objects. We approach the instance segmentation by simply learning the correct embedding space that maps individual instances of objects into distinct clusters that reflect both spatial and semantic information. Unlike previous approaches that require complex pre-processing or post-processing, our implementation is compact and fast with competitive performance, maintaining scalability on large scenes with high resolution voxels. We demonstrate the state-of-the-art performance of our algorithm in the ScanNet 3D instance segmentation benchmark on AP score.