2020/03/18 by David Ifeoluwa Adelani, Michael A. Hedderich, Adelani, David Ifeoluwa +7 · 2 citations
Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Natural Language Processing Techniques #Speech Recognition and Synthesis #Text and Document Classification Technologies #Topic Modeling
paper · pdf · doi:10.48550/arxiv.2003.08370
openalex publication_date 2020/03/18 · openalex created_date 2022/07/26 · openalex updated_date 2026/07/28
The lack of labeled training data has limited the development of natural\nlanguage processing tools, such as named entity recognition, for many languages\nspoken in developing countries. Techniques such as distant and weak supervision\ncan be used to create labeled data in a (semi-) automatic way. Additionally, to\nalleviate some of the negative effects of the errors in automatic annotation,\nnoise-handling methods can be integrated. Pretrained word embeddings are\nanother key component of most neural named entity classifiers. With the advent\nof more complex contextual word embeddings, an interesting trade-off between\nmodel size and performance arises. While these techniques have been shown to\nwork well in high-resource settings, we want to study how they perform in\nlow-resource scenarios. In this work, we perform named entity recognition for\nHausa and Yor `ub 'a, two languages that are widely spoken in several\ndeveloping countries. We evaluate different embedding approaches and show that\ndistant supervision can be successfully leveraged in a realistic low-resource\nscenario where it can more than double a classifier's performance.\n