vix.ing · top · new · best · stats · spec

Seeing through the Human Reporting Bias: Visual Classifiers from Noisy\n Human-Centric Labels

2015/12/22 by Ishan Misra, Misra, Ishan, C. Lawrence Zitnick +5 · 3 citations
Computer Science · #Computer Vision and Pattern Recognition (cs.CV) #Domain Adaptation and Few-Shot Learning #FOS: Computer and information sciences #Machine Learning and Data Classification #Multimodal Machine Learning Applications

paper · pdf · doi:10.48550/arxiv.1512.06974

openalex publication_date 2015/12/22 · openalex created_date 2022/10/06 · openalex updated_date 2026/07/28

Abstract

When human annotators are given a choice about what to label in an image,\nthey apply their own subjective judgments on what to ignore and what to\nmention. We refer to these noisy "human-centric" annotations as exhibiting\nhuman reporting bias. Examples of such annotations include image tags and\nkeywords found on photo sharing sites, or in datasets containing image\ncaptions. In this paper, we use these noisy annotations for learning visually\ncorrect image classifiers. Such annotations do not use consistent vocabulary,\nand miss a significant amount of the information present in an image; however,\nwe demonstrate that the noise in these annotations exhibits structure and can\nbe modeled. We propose an algorithm to decouple the human reporting bias from\nthe correct visually grounded labels. Our results are highly interpretable for\nreporting "what's in the image" versus "what's worth saying." We demonstrate\nthe algorithm's efficacy along a variety of metrics and datasets, including MS\nCOCO and Yahoo Flickr 100M. We show significant improvements over traditional\nalgorithms for both image classification and image captioning, doubling the\nperformance of existing methods in some cases.\n

Cited by

Related