vix.ing · top · new · best · stats

Learning image representations tied to ego-motion

2015/05/08 by Dinesh Jayaraman, Jayaraman, Dinesh, Kristen Grauman +1 · 3 citations
Biochemistry, Genetics and Molecular Biology · Computer Science · Mathematics · #Advanced Vision and Imaging #Artificial Intelligence (cs.AI) #Cell Image Analysis Techniques #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Human Pose and Action Recognition #Machine Learning (stat.ML) #cs.AI #cs.CV #stat.ML

paper · pdf · doi:10.48550/arxiv.1505.02206

Supplementary material appended at end. In ICCV 2015

openalex publication_date 2015/05/08 · arxiv created 2016/03/29 · arxiv updated 2016/03/30 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28

Abstract

Understanding how images of objects and scenes behave in response to specific ego-motions is a crucial aspect of proper visual development, yet existing visual learning methods are conspicuously disconnected from the physical source of their images. We propose to exploit proprioceptive motor signals to provide unsupervised regularization in convolutional neural networks to learn visual representations from egocentric video. Specifically, we enforce that our learned features exhibit equivariance i.e. they respond predictably to transformations associated with distinct ego-motions. With three datasets, we show that our unsupervised feature learning approach significantly outperforms previous approaches on visual recognition and next-best-view prediction tasks. In the most challenging test, we show that features learned from video captured on an autonomous driving platform improve large-scale scene recognition in static images from a disjoint domain.

Citations

Cited by

Related