vix.ing · top · new · best · stats

A Resource-Free Evaluation Metric for Cross-Lingual Word Embeddings Based on Graph Modularity

2019/01/01 by Yoshinari Fujinuma, Jordan Boyd-Graber, Jordan Boyd‐Graber +1 · 21 citations
Computer Science · Mathematics · Physics and Astronomy · #Advanced Graph Neural Networks #Artificial intelligence #Complex Network Analysis Techniques #Computer science #ENCODE #Graph #Mathematics #Metric (unit) #Modularity (biology) #Natural Language Processing Techniques #Natural language processing #Similarity (geometry) #Theoretical computer science #Topic Modeling #Word (group theory) #cs.CL

paper · pdf · doi:10.18653/v1/p19-1489

published in arXiv (Cornell University), 4952-4962 (Cornell University) · Accepted to ACL 2019, camera-ready

arxiv created 2019/06/05 · openalex publication_date 2019/06/05 · openalex created_date 2019/07/30 · arxiv updated 2022/03/24 · openalex updated_date 2026/08/05

Abstract

Cross-lingual word embeddings encode the meaning of words from different languages into a shared low-dimensional space. An important requirement for many downstream tasks is that word similarity should be independent of language - i.e., word vectors within one language should not be more similar to each other than to words in another language. We measure this characteristic using modularity, a network measurement that measures the strength of clusters in a graph. Modularity has a moderate to strong correlation with three downstream tasks, even though modularity is based only on the structure of embeddings and does not require any external resources. We show through experiments that modularity can serve as an intrinsic validation metric to improve unsupervised cross-lingual word embeddings, particularly on distant language pairs in low-resource settings.

Citations