vix.ing · top · new · best · stats

Self-taught Object Localization with Deep Networks

2014/09/13 by Loris Bazzani, Bazzani, Loris, Alessandro Bergamo +5 · 16 citations
Computer Science · #Advanced Image and Video Retrieval Techniques #Advanced Neural Network Applications #Artificial intelligence #Bounding overwatch #Cluster analysis #Cognitive neuroscience of visual object recognition #Computer Vision and Pattern Recognition (cs.CV) #Computer science #Computer vision #Convolutional neural network #Domain Adaptation and Few-Shot Learning #FOS: Computer and information sciences #Ground truth #Image (mathematics) #Key (lock) #Masking (illustration) #Object (grammar) #Object detection #Pattern recognition (psychology) #cs.CV

paper · pdf · doi:10.48550/arxiv.1409.3964

published in arXiv (Cornell University) (Cornell University) · WACV 2016

openalex publication_date 2014/09/13 · arxiv created 2016/02/02 · arxiv updated 2016/02/03 · openalex created_date 2025/10/10 · openalex updated_date 2026/08/06

Abstract

This paper introduces self-taught object localization, a novel approach that leverages deep convolutional networks trained for whole-image recognition to localize objects in images without additional human supervision, i.e., without using any ground-truth bounding boxes for training. The key idea is to analyze the change in the recognition scores when artificially masking out different regions of the image. The masking out of a region that includes the object typically causes a significant drop in recognition score. This idea is embedded into an agglomerative clustering technique that generates self-taught localization hypotheses. Our object localization scheme outperforms existing proposal methods in both precision and recall for small number of subwindow proposals (e.g., on ILSVRC-2012 it produces a relative gain of 23.4% over the state-of-the-art for top-1 hypothesis). Furthermore, our experiments show that the annotations automatically-generated by our method can be used to train object detectors yielding recognition results remarkably close to those obtained by training on manually-annotated bounding boxes.

Citations

Cited by

Related