vix.ing · top · new · best · stats · spec

Exploring Localization for Self-supervised Fine-grained Contrastive Learning

2021/06/30 by Di Wu, Siyuan Li, Wu, Di +5
Computer Science · #Advanced Neural Network Applications #Computer Vision and Pattern Recognition (cs.CV) #Domain Adaptation and Few-Shot Learning #FOS: Computer and information sciences #Visual Attention and Saliency Detection

paper · pdf · doi:10.48550/arxiv.2106.15788

openalex publication_date 2021/06/30 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28

Abstract

Self-supervised contrastive learning has demonstrated great potential in learning visual representations. Despite their success in various downstream tasks such as image classification and object detection, self-supervised pre-training for fine-grained scenarios is not fully explored. We point out that current contrastive methods are prone to memorizing background/foreground texture and therefore have a limitation in localizing the foreground object. Analysis suggests that learning to extract discriminative texture information and localization are equally crucial for fine-grained self-supervised pre-training. Based on our findings, we introduce cross-view saliency alignment (CVSA), a contrastive learning framework that first crops and swaps saliency regions of images as a novel view generation and then guides the model to localize on foreground objects via a cross-view alignment loss. Extensive experiments on both small- and large-scale fine-grained classification benchmarks show that CVSA significantly improves the learned representation.

Related