2021/04/13 by Jae Won Cho, Dong-Jin Kim, Cho, Jae Won +7 · 1 citation
Computer Science · #Advanced Image and Video Retrieval Techniques #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #Domain Adaptation and Few-Shot Learning #FOS: Computer and information sciences #Multimodal Machine Learning Applications
paper · pdf · doi:10.48550/arxiv.2104.05965
openalex publication_date 2021/04/13 · openalex created_date 2022/07/25 · openalex updated_date 2026/07/28
In this work, we address the issues of missing modalities that have arisen\nfrom the Visual Question Answer-Difference prediction task and find a novel\nmethod to solve the task at hand. We address the missing modality-the ground\ntruth answers-that are not present at test time and use a privileged knowledge\ndistillation scheme to deal with the issue of the missing modality. In order to\nefficiently do so, we first introduce a model, the "Big" Teacher, that takes\nthe image/question/answer triplet as its input and outperforms the baseline,\nthen use a combination of models to distill knowledge to a target network\n(student) that only takes the image/question pair as its inputs. We experiment\nour models on the VizWiz and VQA-V2 Answer Difference datasets and show through\nextensive experimentation and ablation the performances of our method and a\ndiverse possibility for future research.\n