2017/06/30 by Sebastian Ruder, Ivan Vulić, Anders Søgaard · 256 citations
Computer Science · #Embedding #ICT in Developing Communities #Key (lock) #Meaning (existential) #Natural Language Processing Techniques #Theme (computing) #Topic Modeling #Word (group theory) #Word embedding #cs.CL #cs.LG
paper · pdf · doi:10.1613/jair.1.11640
published in Journal of Artificial Intelligence Research 65, 569-631 (AI Access Foundation) · Published in Journal of Artificial Intelligence Research
openalex created_date 2017/12/04 · openalex publication_date 2019/08/12 · arxiv created 2019/10/06 · arxiv updated 2019/10/08 · openalex updated_date 2026/08/05
Cross-lingual representations of words enable us to reason about word meaning in multilingual contexts and are a key facilitator of cross-lingual transfer when developing natural language processing models for low-resource languages. In this survey, we provide a comprehensive typology of cross-lingual word embedding models. We compare their data requirements and objective functions. The recurring theme of the survey is that many of the models presented in the literature optimize for the same objectives, and that seemingly different models are often equivalent, modulo optimization strategies, hyper-parameters, and such. We also discuss the different ways cross-lingual word embeddings are evaluated, as well as future challenges and research horizons.