vix.ing · top · new · best · stats · spec

Lexical Normalization for Code-switched Data and its Effect on\n POS-tagging

2020/06/01 by Rob van der Goot, van der Goot, Rob, Özlem Çetinoğlu +1
Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Natural Language Processing Techniques #Speech Recognition and Synthesis #Topic Modeling

paper · pdf · doi:10.48550/arxiv.2006.01175

openalex publication_date 2020/06/01 · openalex created_date 2022/07/26 · openalex updated_date 2026/07/28

Abstract

Lexical normalization, the translation of non-canonical data to standard\nlanguage, has shown to improve the performance of manynatural language\nprocessing tasks on social media. Yet, using multiple languages in one\nutterance, also called code-switching (CS), is frequently overlooked by these\nnormalization systems, despite its common use in social media. In this paper,\nwe propose three normalization models specifically designed to handle\ncode-switched data which we evaluate for two language pairs: Indonesian-English\n(Id-En) and Turkish-German (Tr-De). For the latter, we introduce novel\nnormalization layers and their corresponding language ID and POS tags for the\ndataset, and evaluate the downstream effect of normalization on POS tagging.\nResults show that our CS-tailored normalization models outperform Id-En state\nof the art and Tr-De monolingual models, and lead to 5.4% relative performance\nincrease for POS tagging as compared to unnormalized input.\n

Cited by

Related