2017/05/11 by Desmond Elliott, Elliott, Desmond, Ákos Kádár +1 · 7 citations
Computer Science · #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Multimodal Machine Learning Applications #Natural Language Processing Techniques #Topic Modeling #cs.CL #cs.CV
paper · pdf · doi:10.48550/arxiv.1705.04350
Clarified main contributions, minor correction to Equation 8, additional comparisons in Table 2, added more related work
openalex publication_date 2017/05/11 · arxiv created 2017/07/07 · arxiv updated 2017/07/10 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
We decompose multimodal translation into two sub-tasks: learning to translate and learning visually grounded representations. In a multitask learning framework, translations are learned in an attention-based encoder-decoder, and grounded representations are learned through image representation prediction. Our approach improves translation performance compared to the state of the art on the Multi30K dataset. Furthermore, it is equally effective if we train the image prediction task on the external MS COCO dataset, and we find improvements if we train the translation model on the external News Commentary parallel text.