2019/10/28 by Nasser Zalmout, Zalmout, Nasser, Nizar Habash +1
Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Natural Language Processing Techniques #Speech Recognition and Synthesis #Topic Modeling
paper · pdf · doi:10.48550/arxiv.1910.12702
openalex publication_date 2019/10/28 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Morphological tagging is challenging for morphologically rich languages due\nto the large target space and the need for more training data to minimize model\nsparsity. Dialectal variants of morphologically rich languages suffer more as\nthey tend to be more noisy and have less resources. In this paper we explore\nthe use of multitask learning and adversarial training to address morphological\nrichness and dialectal variations in the context of full morphological tagging.\nWe use multitask learning for joint morphological modeling for the features\nwithin two dialects, and as a knowledge-transfer scheme for cross-dialectal\nmodeling. We use adversarial training to learn dialect invariant features that\ncan help the knowledge-transfer scheme from the high to low-resource variants.\nWe work with two dialectal variants: Modern Standard Arabic (high-resource\n"dialect") and Egyptian Arabic (low-resource dialect) as a case study. Our\nmodels achieve state-of-the-art results for both. Furthermore, adversarial\ntraining provides more significant improvement when using smaller training\ndatasets in particular.\n