2021/08/29 by Injy Hamed, Pavel Denisov, Hamed, Injy +10 · 1 citation
Computer Science · Psychology · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Natural Language Processing Techniques #Phonetics and Phonology Research #Speech Recognition and Synthesis #Speech and dialogue systems
paper · pdf · doi:10.48550/arxiv.2108.12881
openalex publication_date 2021/08/29 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Code-switching (CS), defined as the mixing of languages in conversations, has\nbecome a worldwide phenomenon. The prevalence of CS has been recently met with\na growing demand and interest to build CS ASR systems. In this paper, we\npresent our work on code-switched Egyptian Arabic-English automatic speech\nrecognition (ASR). We first contribute in filling the huge gap in resources by\ncollecting, analyzing and publishing our spontaneous CS Egyptian Arabic-English\nspeech corpus. We build our ASR systems using DNN-based hybrid and\nTransformer-based end-to-end models. In this paper, we present a thorough\ncomparison between both approaches under the setting of a low-resource,\northographically unstandardized, and morphologically rich language pair. We\nshow that while both systems give comparable overall recognition results, each\nsystem provides complementary sets of strength points. We show that recognition\ncan be improved by combining the outputs of both systems. We propose several\neffective system combination approaches, where hypotheses of both systems are\nmerged on sentence- and word-levels. Our approaches result in overall WER\nrelative improvement of 4.7%, over a baseline performance of 32.1% WER. In the\ncase of intra-sentential CS sentences, we achieve WER relative improvement of\n4.8%. Our best performing system achieves 30.6% WER on ArzEn test set.\n