vix.ing · top · new · best · stats · spec

Learning Conditioned Graph Structures for Interpretable Visual Question\n Answering

2018/06/19 by Will Norcliffe-Brown, Norcliffe-Brown, Will, Efstathios Vafeias +3 · 4 citations
Computer Science · #Advanced Image and Video Retrieval Techniques #Computer Vision and Pattern Recognition (cs.CV) #Domain Adaptation and Few-Shot Learning #FOS: Computer and information sciences #Multimodal Machine Learning Applications

paper · pdf · doi:10.48550/arxiv.1806.07243

openalex publication_date 2018/06/19 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28

Abstract

Visual Question answering is a challenging problem requiring a combination of\nconcepts from Computer Vision and Natural Language Processing. Most existing\napproaches use a two streams strategy, computing image and question features\nthat are consequently merged using a variety of techniques. Nonetheless, very\nfew rely on higher level image representations, which can capture semantic and\nspatial relationships. In this paper, we propose a novel graph-based approach\nfor Visual Question Answering. Our method combines a graph learner module,\nwhich learns a question specific graph representation of the input image, with\nthe recent concept of graph convolutions, aiming to learn image representations\nthat capture question specific interactions. We test our approach on the VQA v2\ndataset using a simple baseline architecture enhanced by the proposed graph\nlearner module. We obtain promising results with 66.18% accuracy and\ndemonstrate the interpretability of the proposed method. Code can be found at\ngithub.com/aimbrain/vqa-project.\n

Cited by

Related