vix.ing · top · new · best · stats · spec

Hindi Visual Genome: A Dataset for Multimodal English-to-Hindi Machine\n Translation

2019/07/21 by Shantipriya Parida, Parida, Shantipriya, Ondřej Bojar +3 · 1 citation
Computer Science · Biochemistry, Genetics and Molecular Biology · #Multimodal Machine Learning Applications #Advanced Image and Video Retrieval Techniques #Genomics and Phylogenetic Studies

paper · pdf · doi:10.48550/arxiv.1907.08948

Abstract

Visual Genome is a dataset connecting structured image information with\nEnglish language. We present ``Hindi Visual Genome'', a multimodal dataset\nconsisting of text and images suitable for English-Hindi multimodal machine\ntranslation task and multimodal research. We have selected short English\nsegments (captions) from Visual Genome along with associated images and\nautomatically translated them to Hindi with manual post-editing which took the\nassociated images into account. We prepared a set of 31525 segments,\naccompanied by a challenge test set of 1400 segments. This challenge test set\nwas created by searching for (particularly) ambiguous English words based on\nthe embedding similarity and manually selecting those where the image helps to\nresolve the ambiguity.\n Our dataset is the first for multimodal English-Hindi machine translation,\nfreely available for non-commercial research purposes. Our Hindi version of\nVisual Genome also allows to create Hindi image labelers or other practical\ntools.\n Hindi Visual Genome also serves in Workshop on Asian Translation (WAT) 2019\nMulti-Modal Translation Task.\n

Cited by

Related