2020/06/12 by Siddhesh Khandelwal, Raghav Goyal, Khandelwal, Siddhesh +3
Computer Science · #Advanced Image and Video Retrieval Techniques #Advanced Neural Network Applications #Computer Vision and Pattern Recognition (cs.CV) #Domain Adaptation and Few-Shot Learning #FOS: Computer and information sciences
paper · pdf · doi:10.48550/arxiv.2006.07502
openalex publication_date 2020/06/12 · openalex created_date 2022/07/26 · openalex updated_date 2026/07/28
Methods for object detection and segmentation rely on large scale\ninstance-level annotations for training, which are difficult and time-consuming\nto collect. Efforts to alleviate this look at varying degrees and quality of\nsupervision. Weakly-supervised approaches draw on image-level labels to build\ndetectors/segmentors, while zero/few-shot methods assume abundant\ninstance-level data for a set of base classes, and none to a few examples for\nnovel classes. This taxonomy has largely siloed algorithmic designs. In this\nwork, we aim to bridge this divide by proposing an intuitive and unified\nsemi-supervised model that is applicable to a range of supervision: from zero\nto a few instance-level samples per novel class. For base classes, our model\nlearns a mapping from weakly-supervised to fully-supervised\ndetectors/segmentors. By learning and leveraging visual and lingual\nsimilarities between the novel and base classes, we transfer those mappings to\nobtain detectors/segmentors for novel classes; refining them with a few novel\nclass instance-level annotated samples, if available. The overall model is\nend-to-end trainable and highly flexible. Through extensive experiments on\nMS-COCO and Pascal VOC benchmark datasets we show improved performance in a\nvariety of settings.\n