vix.ing · top · new · best · stats

Clustering Comparable Corpora of Russian and Ukrainian Academic Texts: Word Embeddings and Semantic Fingerprints

2016/04/18 by Andrey Kutuzov, Михаил Копотев, Mikhail Kopotev +6
Computer Science · #Advanced Text Analysis Techniques #Computation and Language (cs.CL) #FOS: Computer and information sciences #Natural Language Processing Techniques #Topic Modeling #cs.CL

paper · pdf · doi:10.48550/arxiv.1604.05372

To be presented at 9th Workshop on Building and Using Comparable Corpora, co-located with LREC-2016 (https://comparable.limsi.fr/bucc2016/)

arxiv created 2016/04/18 · openalex publication_date 2016/04/18 · arxiv updated 2016/04/20 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28

Abstract

We present our experience in applying distributional semantics (neural word embeddings) to the problem of representing and clustering documents in a bilingual comparable corpus. Our data is a collection of Russian and Ukrainian academic texts, for which topics are their academic fields. In order to build language-independent semantic representations of these documents, we train neural distributional models on monolingual corpora and learn the optimal linear transformation of vectors from one language to another. The resulting vectors are then used to produce `semantic fingerprints' of documents, serving as input to a clustering algorithm. The presented method is compared to several baselines including `orthographic translation' with Levenshtein edit distance and outperforms them by a large margin. We also show that language-independent `semantic fingerprints' are superior to multi-lingual clustering algorithms proposed in the previous work, at the same time requiring less linguistic resources.

Citations

Related