vix.ing · top · new · best · stats

Attention on Attention: Architectures for Visual Question Answering (VQA)

2018/03/21 by Jasdeep Singh, Singh, Jasdeep, Vincent Ying +3 · 1 citation
Computer Science · #68Txx #Advanced Image and Video Retrieval Techniques #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #Domain Adaptation and Few-Shot Learning #FOS: Computer and information sciences #Multimodal Machine Learning Applications #cs.AI #cs.CL #cs.CV #msc:68Txx

paper · pdf · doi:10.48550/arxiv.1803.07724

Visual Question Answering Project

arxiv created 2018/03/21 · openalex publication_date 2018/03/21 · arxiv updated 2018/03/22 · openalex created_date 2018/03/29 · openalex updated_date 2026/07/28

Abstract

Visual Question Answering (VQA) is an increasingly popular topic in deep learning research, requiring coordination of natural language processing and computer vision modules into a single architecture. We build upon the model which placed first in the VQA Challenge by developing thirteen new attention mechanisms and introducing a simplified classifier. We performed 300 GPU hours of extensive hyperparameter and architecture searches and were able to achieve an evaluation score of 64.78%, outperforming the existing state-of-the-art single model's validation score of 63.15%.

Citations

Cited by

Related