vix.ing · top · new · best · stats

Unsupervised Object-Level Representation Learning from Scene Images

2021/06/22 by Jiahao Xie, Xie, Jiahao, Xiaohang Zhan +8 · 11 citations
Computer Science · #Advanced Neural Network Applications #Artificial intelligence #Artificial neural network #Cognitive neuroscience of visual object recognition #Computer Vision and Pattern Recognition (cs.CV) #Computer science #Computer vision #Domain Adaptation and Few-Shot Learning #FOS: Computer and information sciences #Feature learning #Learning object #Leverage (statistics) #Machine learning #Multimodal Machine Learning Applications #Object (grammar) #Pattern recognition (psychology) #Representation (politics) #Semi-supervised learning #Supervised learning #Unsupervised learning #cs.CV

paper · pdf · doi:10.48550/arxiv.2106.11952

published in arXiv (Cornell University) (Cornell University) · NeurIPS 2021. Project page: https://www.mmlab-ntu.com/project/orl/ Code: https://github.com/Jiahao000/ORL

openalex publication_date 2021/06/22 · arxiv created 2021/12/03 · arxiv updated 2021/12/06 · openalex created_date 2025/10/10 · openalex updated_date 2026/08/05

Abstract

Contrastive self-supervised learning has largely narrowed the gap to supervised pre-training on ImageNet. However, its success highly relies on the object-centric priors of ImageNet, i.e., different augmented views of the same image correspond to the same object. Such a heavily curated constraint becomes immediately infeasible when pre-trained on more complex scene images with many objects. To overcome this limitation, we introduce Object-level Representation Learning (ORL), a new self-supervised learning framework towards scene images. Our key insight is to leverage image-level self-supervised pre-training as the prior to discover object-level semantic correspondence, thus realizing object-level representation learning from scene images. Extensive experiments on COCO show that ORL significantly improves the performance of self-supervised learning on scene images, even surpassing supervised ImageNet pre-training on several downstream tasks. Furthermore, ORL improves the downstream performance when more unlabeled scene images are available, demonstrating its great potential of harnessing unlabeled data in the wild. We hope our approach can motivate future research on more general-purpose unsupervised representation learning from scene data.

Citations

Cited by

Related