vix.ing · top · new · best · stats · spec

Introducing two Vietnamese Datasets for Evaluating Semantic Models of\n (Dis-)Similarity and Relatedness

2018/04/15 by Kim-Anh Nguyen, Nguyen, Kim Anh, Sabine Schulte im Walde +3 · 1 citation
Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Natural Language Processing Techniques #Topic Modeling

paper · pdf · doi:10.48550/arxiv.1804.05388

openalex publication_date 2018/04/15 · openalex created_date 2022/10/01 · openalex updated_date 2026/07/28

Abstract

We present two novel datasets for the low-resource language Vietnamese to\nassess models of semantic similarity: ViCon comprises pairs of synonyms and\nantonyms across word classes, thus offering data to distinguish between\nsimilarity and dissimilarity. ViSim-400 provides degrees of similarity across\nfive semantic relations, as rated by human judges. The two datasets are\nverified through standard co-occurrence and neural network models, showing\nresults comparable to the respective English datasets.\n

Cited by

Related