vix.ing · top · new · best · stats

CLUSE: Cross-Lingual Unsupervised Sense Embeddings

2018/09/15 by Ta-Chung Chi, Chi, Ta-Chung, Yun-Nung Chen +1
Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Natural Language Processing Techniques #Speech and dialogue systems #Topic Modeling #cs.CL

paper · pdf · doi:10.48550/arxiv.1809.05694

11 pages, accepted by EMNLP 2018

openalex publication_date 2018/09/15 · arxiv created 2018/10/21 · arxiv updated 2018/10/23 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28

Abstract

This paper proposes a modularized sense induction and representation learning model that jointly learns bilingual sense embeddings that align well in the vector space, where the cross-lingual signal in the English-Chinese parallel corpus is exploited to capture the collocation and distributed characteristics in the language pair. The model is evaluated on the Stanford Contextual Word Similarity (SCWS) dataset to ensure the quality of monolingual sense embeddings. In addition, we introduce Bilingual Contextual Word Similarity (BCWS), a large and high-quality dataset for evaluating cross-lingual sense embeddings, which is the first attempt of measuring whether the learned embeddings are indeed aligned well in the vector space. The proposed approach shows the superior quality of sense embeddings evaluated in both monolingual and bilingual spaces.

Citations

Related