vix.ing · top · new · best · stats · spec

Attribute Diversity Determines the Systematicity Gap in VQA

2023/11/15 by Ian Berlot-Attwell, Berlot-Attwell, Ian, Kumar Krishna Agrawal +7 · 1 voice
Computer Science · #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Machine Learning (cs.LG) #cs.CL #cs.CV #cs.LG

paper · pdf · doi:10.48550/arxiv.2311.08695

arxiv published 2023/11/15 · arxiv updated 2024/10/04

Abstract

Although modern neural networks often generalize to new combinations of familiar concepts, the conditions that enable such compositionality have long been an open question. In this work, we study the systematicity gap in visual question answering: the performance difference between reasoning on previously seen and unseen combinations of object attributes. To test, we introduce a novel diagnostic dataset, CLEVR-HOPE. We find that the systematicity gap is not reduced by increasing the quantity of training data, but is reduced by increasing the diversity of training data. In particular, our experiments suggest that the more distinct attribute type combinations are seen during training, the more systematic we can expect the resulting model to be.

Discussions

Related