vix.ing · top · new · best · stats · spec

Learning to Disambiguate by Asking Discriminative Questions

2017/08/09 by Yining Li, Chen Huang, Li, Yining +6
Computer Science · Mathematics · #Advanced Image and Video Retrieval Techniques #Artificial intelligence #Benchmarking #Closed captioning #Complement (music) #Computer Vision and Pattern Recognition (cs.CV) #Computer science #Discriminative model #Domain Adaptation and Few-Shot Learning #FOS: Computer and information sciences #Image (mathematics) #Information retrieval #Machine learning #Mathematics #Multimodal Machine Learning Applications #Natural language processing #Pattern recognition (psychology) #Question answering #Tuple #Visualization #cs.CV

paper · pdf · doi:10.48550/arxiv.1708.02760

14 pages, 12 figures, ICCV2017

arxiv created 2017/08/09 · openalex publication_date 2017/08/09 · arxiv updated 2017/08/10 · openalex created_date 2025/10/10 · openalex updated_date 2026/08/05

Abstract

The ability to ask questions is a powerful tool to gather information in order to learn about the world and resolve ambiguities. In this paper, we explore a novel problem of generating discriminative questions to help disambiguate visual instances. Our work can be seen as a complement and new extension to the rich research studies on image captioning and question answering. We introduce the first large-scale dataset with over 10,000 carefully annotated images-question tuples to facilitate benchmarking. In particular, each tuple consists of a pair of images and 4.6 discriminative questions (as positive samples) and 5.9 non-discriminative questions (as negative samples) on average. In addition, we present an effective method for visual discriminative question generation. The method can be trained in a weakly supervised manner without discriminative images-question tuples but just existing visual question answering datasets. Promising results are shown against representative baselines through quantitative evaluations and user studies.

Citations

Related