vix.ing · top · new · best · stats · spec

Unsupervised Morphological Expansion of Small Datasets for Improving Word Embeddings

2017/11/15 by Syed Sarfaraz Akhtar, Akhtar, Syed Sarfaraz, Arihant Gupta +7
Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Natural Language Processing Techniques #Text and Document Classification Technologies #Topic Modeling

paper · pdf · doi:10.48550/arxiv.1711.05678

openalex publication_date 2017/11/15 · openalex created_date 2017/12/04 · openalex updated_date 2026/07/28

Abstract

We present a language independent, unsupervised method for building word embeddings using morphological expansion of text. Our model handles the problem of data sparsity and yields improved word embeddings by relying on training word embeddings on artificially generated sentences. We evaluate our method using small sized training sets on eleven test sets for the word similarity task across seven languages. Further, for English, we evaluated the impacts of our approach using a large training set on three standard test sets. Our method improved results across all languages.

Related