2018/03/21 by Jasdeep Singh, Singh, Jasdeep, Vincent Ying +3 · 1 citation
Computer Science · #68Txx #Advanced Image and Video Retrieval Techniques #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #Domain Adaptation and Few-Shot Learning #FOS: Computer and information sciences #Multimodal Machine Learning Applications #cs.AI #cs.CL #cs.CV #msc:68Txx
paper · pdf · doi:10.48550/arxiv.1803.07724
Visual Question Answering Project
arxiv created 2018/03/21 · openalex publication_date 2018/03/21 · arxiv updated 2018/03/22 · openalex created_date 2018/03/29 · openalex updated_date 2026/07/28
Visual Question Answering (VQA) is an increasingly popular topic in deep learning research, requiring coordination of natural language processing and computer vision modules into a single architecture. We build upon the model which placed first in the VQA Challenge by developing thirteen new attention mechanisms and introducing a simplified classifier. We performed 300 GPU hours of extensive hyperparameter and architecture searches and were able to achieve an evaluation score of 64.78%, outperforming the existing state-of-the-art single model's validation score of 63.15%.