2021/05/08 by Maxime Kayser, Oana-Maria Camburu, Kayser, Maxime +11 · 5 citations
Computer Science · #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #Explainable Artificial Intelligence (XAI) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Multimodal Machine Learning Applications #Topic Modeling
paper · pdf · doi:10.48550/arxiv.2105.03761
openalex publication_date 2021/05/08 · openalex created_date 2022/07/25 · openalex updated_date 2026/07/28
Recently, there has been an increasing number of efforts to introduce models\ncapable of generating natural language explanations (NLEs) for their\npredictions on vision-language (VL) tasks. Such models are appealing, because\nthey can provide human-friendly and comprehensive explanations. However, there\nis a lack of comparison between existing methods, which is due to a lack of\nre-usable evaluation frameworks and a scarcity of datasets. In this work, we\nintroduce e-ViL and e-SNLI-VE. e-ViL is a benchmark for explainable\nvision-language tasks that establishes a unified evaluation framework and\nprovides the first comprehensive comparison of existing approaches that\ngenerate NLEs for VL tasks. It spans four models and three datasets and both\nautomatic metrics and human evaluation are used to assess model-generated\nexplanations. e-SNLI-VE is currently the largest existing VL dataset with NLEs\n(over 430k instances). We also propose a new model that combines UNITER, which\nlearns joint embeddings of images and text, and GPT-2, a pre-trained language\nmodel that is well-suited for text generation. It surpasses the previous state\nof the art by a large margin across all datasets. Code and data are available\nhere: https://github.com/maximek3/e-ViL.\n