2018/01/23 by Martin Thoma, Thoma, Martin · 2 citations
Computer Science · #Natural Language Processing Techniques #Speech Recognition and Synthesis #Topic Modeling #cs.CL #cs.CV
paper · pdf · doi:10.48550/arxiv.1801.07779
{"pages": 12, "figures": 4, "language": "English", "author-ORCiD": ["https://orcid.org/0000-0002-6517-1690"]}
arxiv created 2018/01/23 · arxiv updated 2018/01/25
This paper describes the WiLI-2018 benchmark dataset for monolingual written natural language identification. WiLI-2018 is a publicly available, free of charge dataset of short text extracts from Wikipedia. It contains 1000 paragraphs of 235 languages, totaling in 23500 paragraphs. WiLI is a classification dataset: Given an unknown paragraph written in one dominant language, it has to be decided which language it is.