2016/09/27 by Magnus Sahlgren, Sahlgren, Magnus, Alessandro Lenci +1 · 1 citation
Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning in Healthcare #Natural Language Processing Techniques #Topic Modeling
paper · pdf · doi:10.48550/arxiv.1609.08293
openalex publication_date 2016/09/27 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
This paper investigates the effects of data size and frequency range on distributional semantic models. We compare the performance of a number of representative models for several test settings over data of varying sizes, and over test items of various frequency. Our results show that neural network-based models underperform when the data is small, and that the most reliable model over data of varying sizes and frequency ranges is the inverted factorized model.