2018/04/15 by Kim-Anh Nguyen, Nguyen, Kim Anh, Sabine Schulte im Walde +3 · 1 citation
Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Natural Language Processing Techniques #Topic Modeling
paper · pdf · doi:10.48550/arxiv.1804.05388
openalex publication_date 2018/04/15 · openalex created_date 2022/10/01 · openalex updated_date 2026/07/28
We present two novel datasets for the low-resource language Vietnamese to\nassess models of semantic similarity: ViCon comprises pairs of synonyms and\nantonyms across word classes, thus offering data to distinguish between\nsimilarity and dissimilarity. ViSim-400 provides degrees of similarity across\nfive semantic relations, as rated by human judges. The two datasets are\nverified through standard co-occurrence and neural network models, showing\nresults comparable to the respective English datasets.\n