2020/09/30 by Kiet Van Nguyen, Duc-Vu Nguyen, Van Nguyen, Kiet +5 · 7 citations
Computer Science · Mathematics · #Artificial intelligence #Benchmark (surveying) #Comprehension #Computation and Language (cs.CL) #Computer science #FOS: Computer and information sciences #Geography #Linguistics #Machine translation #Matching (statistics) #Mathematics #Multimodal Machine Learning Applications #Natural Language Processing Techniques #Natural language processing #Question answering #Reading (process) #Sentence #Task (project management) #Topic Modeling #Vietnamese #cs.CL
paper · pdf · doi:10.48550/arxiv.2009.14725
published in arXiv (Cornell University) (Cornell University) · Accepted by The 28th International Conference on Computational Linguistics (COLING 2020)
openalex publication_date 2020/09/30 · arxiv created 2020/11/07 · arxiv updated 2020/11/10 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Over 97 million people speak Vietnamese as their native language in the world. However, there are few research studies on machine reading comprehension (MRC) for Vietnamese, the task of understanding a text and answering questions related to it. Due to the lack of benchmark datasets for Vietnamese, we present the Vietnamese Question Answering Dataset (UIT-ViQuAD), a new dataset for the low-resource language as Vietnamese to evaluate MRC models. This dataset comprises over 23,000 human-generated question-answer pairs based on 5,109 passages of 174 Vietnamese articles from Wikipedia. In particular, we propose a new process of dataset creation for Vietnamese MRC. Our in-depth analyses illustrate that our dataset requires abilities beyond simple reasoning like word matching and demands single-sentence and multiple-sentence inferences. Besides, we conduct experiments on state-of-the-art MRC methods for English and Chinese as the first experimental models on UIT-ViQuAD. We also estimate human performance on the dataset and compare it to the experimental results of powerful machine learning models. As a result, the substantial differences between human performance and the best model performance on the dataset indicate that improvements can be made on UIT-ViQuAD in future research. Our dataset is freely available on our website to encourage the research community to overcome challenges in Vietnamese MRC.