vix.ing · top · new · best · stats · spec

Improving Yor\`ub'a Diacritic Restoration

2020/03/23 by Iroro Orife, David Ifeoluwa Adelani, Orife, Iroro +11
Agricultural and Biological Sciences · Computer Science · #Agriculture and Rural Development Research #Botany and Geology in Latin America and Caribbean #Computation and Language (cs.CL) #FOS: Computer and information sciences #Natural Language Processing Techniques #Speech Recognition and Synthesis #Topic Modeling

paper · pdf · doi:10.48550/arxiv.2003.10564

openalex publication_date 2020/03/23 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28

Abstract

Yor `ub 'a is a widely spoken West African language with a writing system\nrich in orthographic and tonal diacritics. They provide morphological\ninformation, are crucial for lexical disambiguation, pronunciation and are\nvital for any computational Speech or Natural Language Processing tasks.\nHowever diacritic marks are commonly excluded from electronic texts due to\nlimited device and application support as well as general education on proper\nusage. We report on recent efforts at dataset cultivation. By aggregating and\nimproving disparate texts from the web and various personal libraries, we were\nable to significantly grow our clean Yor `ub 'a dataset from a majority\nBibilical text corpora with three sources to millions of tokens from over a\ndozen sources. We evaluate updated diacritic restoration models on a new,\ngeneral purpose, public-domain Yor `ub 'a evaluation dataset of modern\njournalistic news text, selected to be multi-purpose and reflecting\ncontemporary usage. All pre-trained models, datasets and source-code have been\nreleased as an open-source project to advance efforts on Yor `ub 'a language\ntechnology.\n

Related