vix.ing · top · new · best · stats

Attend, Infer, Repeat: Fast Scene Understanding with Generative Models

2016/03/28 by S. M. Ali Eslami, Eslami, S. M. Ali, Nicolas Heess +11 · 1 voice · 10 citations
Computer Science · #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Machine Learning (cs.LG) #cs.CV #cs.LG

paper · pdf · doi:10.48550/arxiv.1603.08575

arxiv published 2016/03/28 · arxiv created 2016/08/12 · arxiv updated 2016/08/15

Abstract

We present a framework for efficient inference in structured image models that explicitly reason about objects. We achieve this by performing probabilistic inference using a recurrent neural network that attends to scene elements and processes them one at a time. Crucially, the model itself learns to choose the appropriate number of inference steps. We use this scheme to learn to perform inference in partially specified 2D models (variable-sized variational auto-encoders) and fully specified 3D models (probabilistic renderers). We show that such models learn to identify multiple objects - counting, locating and classifying the elements of a scene - without any supervision, e.g., decomposing 3D images with various numbers of objects in a single forward pass of a neural network. We further show that the networks produce accurate inferences when compared to supervised counterparts, and that their structure leads to improved generalization.

Cited by

Discussions

Related