2012/06/27 by Yun Jiang, Jiang Yun, Marcus Lim +4 · 5 citations
Computer Science · Mathematics · #Advanced Image and Video Retrieval Techniques #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Image Processing and 3D Reconstruction #Image Retrieval and Classification Techniques #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Robotics (cs.RO) #cs.CV #cs.LG #cs.RO #stat.ML
paper · pdf · doi:10.48550/arxiv.1206.6462
Appears in Proceedings of the 29th International Conference on Machine Learning (ICML 2012)
arxiv created 2012/06/27 · openalex publication_date 2012/06/27 · arxiv updated 2012/07/02 · openalex created_date 2016/06/24 · openalex updated_date 2026/07/28
We consider the problem of learning object arrangements in a 3D scene. The key idea here is to learn how objects relate to human poses based on their affordances, ease of use and reachability. In contrast to modeling object-object relationships, modeling human-object relationships scales linearly in the number of objects. We design appropriate density functions based on 3D spatial features to capture this. We learn the distribution of human poses in a scene using a variant of the Dirichlet process mixture model that allows sharing of the density function parameters across the same object types. Then we can reason about arrangements of the objects in the room based on these meaningful human poses. In our extensive experiments on 20 different rooms with a total of 47 objects, our algorithm predicted correct placements with an average error of 1.6 meters from ground truth. In arranging five real scenes, it received a score of 4.3/5 compared to 3.7 for the best baseline method.