2020/07/24 by Yilun Du, Du, Yilun, Kevin A. Smith +7 · 1 citation
Computer Science · Engineering · #3D Shape Modeling and Analysis #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Human Pose and Action Recognition #Image Processing and 3D Reconstruction #Machine Learning (cs.LG)
paper · pdf · doi:10.48550/arxiv.2007.12348
openalex publication_date 2020/07/24 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
We study the problem of unsupervised physical object discovery. While existing frameworks aim to decompose scenes into 2D segments based off each object's appearance, we explore how physics, especially object interactions, facilitates disentangling of 3D geometry and position of objects from video, in an unsupervised manner. Drawing inspiration from developmental psychology, our Physical Object Discovery Network (POD-Net) uses both multi-scale pixel cues and physical motion cues to accurately segment observable and partially occluded objects of varying sizes, and infer properties of those objects. Our model reliably segments objects on both synthetic and real scenes. The discovered object properties can also be used to reason about physical events.