2017/01/10 by Danfei Xu, Xu, Danfei, Yuke Zhu +6 · 102 citations
Computer Science · #Advanced Image and Video Retrieval Techniques #Artificial intelligence #Computer science #Computer vision #Graph #Graphical model #Image Retrieval and Classification Techniques #Inference #Machine learning #Message passing #Multimodal Machine Learning Applications #Pattern recognition (psychology) #Representation (politics) #Scene graph #Theoretical computer science #cs.CV
paper · pdf · doi:10.48550/arxiv.1701.02426
published in arXiv (Cornell University) (Cornell University) · CVPR 2017
openalex publication_date 2017/01/10 · arxiv created 2017/04/12 · arxiv updated 2017/04/13 · openalex created_date 2025/10/10 · openalex updated_date 2026/08/04
Understanding a visual scene goes beyond recognizing individual objects in isolation. Relationships between objects also constitute rich semantic information about the scene. In this work, we explicitly model the objects and their relationships using scene graphs, a visually-grounded graphical structure of an image. We propose a novel end-to-end model that generates such structured scene representation from an input image. The model solves the scene graph inference problem using standard RNNs and learns to iteratively improves its predictions via message passing. Our joint inference model can take advantage of contextual cues to make better predictions on objects and their relationships. The experiments show that our model significantly outperforms previous methods for generating scene graphs using Visual Genome dataset and inferring support relations with NYU Depth v2 dataset.