vix.ing · top · new · best · stats

Spatially Invariant Unsupervised 3D Object-Centric Learning and Scene Decomposition

2021/06/10 by Tianyu Wang, Miaomiao Liu, Wang, Tianyu +3
Computer Science · #Advanced Image and Video Retrieval Techniques #Advanced Neural Network Applications #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Human Pose and Action Recognition #cs.CV

paper · pdf · doi:10.48550/arxiv.2106.05607

openalex publication_date 2021/06/10 · arxiv created 2022/07/17 · arxiv updated 2022/07/19 · openalex created_date 2022/07/22 · openalex updated_date 2026/07/28

Abstract

We tackle the problem of object-centric learning on point clouds, which is crucial for high-level relational reasoning and scalable machine intelligence. In particular, we introduce a framework, SPAIR3D, to factorize a 3D point cloud into a spatial mixture model where each component corresponds to one object. To model the spatial mixture model on point clouds, we derive the Chamfer Mixture Loss, which fits naturally into our variational training pipeline. Moreover, we adopt an object-specification scheme that describes each object's location relative to its local voxel grid cell. Such a scheme allows SPAIR3D to model scenes with an arbitrary number of objects. We evaluate our method on the task of unsupervised scene decomposition. Experimental results demonstrate that SPAIR3D has strong scalability and is capable of detecting and segmenting an unknown number of objects from a point cloud in an unsupervised manner.

Related