2015/09/24 by Ivan Vulić, Vulić, Ivan, Marie‐Francine Moens +1
Computer Science · #Advanced Text Analysis Techniques #Computation and Language (cs.CL) #FOS: Computer and information sciences #Natural Language Processing Techniques #Topic Modeling
paper · pdf · doi:10.48550/arxiv.1509.07308
openalex publication_date 2015/09/24 · openalex created_date 2022/10/01 · openalex updated_date 2026/07/28
We propose a new model for learning bilingual word representations from\nnon-parallel document-aligned data. Following the recent advances in word\nrepresentation learning, our model learns dense real-valued word vectors, that\nis, bilingual word embeddings (BWEs). Unlike prior work on inducing BWEs which\nheavily relied on parallel sentence-aligned corpora and/or readily available\ntranslation resources such as dictionaries, the article reveals that BWEs may\nbe learned solely on the basis of document-aligned comparable data without any\nadditional lexical resources nor syntactic information. We present a comparison\nof our approach with previous state-of-the-art models for learning bilingual\nword representations from comparable data that rely on the framework of\nmultilingual probabilistic topic modeling (MuPTM), as well as with\ndistributional local context-counting models. We demonstrate the utility of the\ninduced BWEs in two semantic tasks: (1) bilingual lexicon extraction, (2)\nsuggesting word translations in context for polysemous words. Our simple yet\neffective BWE-based models significantly outperform the MuPTM-based and\ncontext-counting representation models from comparable data as well as prior\nBWE-based models, and acquire the best reported results on both tasks for all\nthree tested language pairs.\n