2020/06/09 by Gurunath Reddy Madhumani, Sanket Shah, Madhumani, Gurunath Reddy +7
Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #Domain Adaptation and Few-Shot Learning #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
paper · pdf · doi:10.48550/arxiv.2006.05257
openalex publication_date 2020/06/09 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Recognizing code-switched speech is challenging for Automatic Speech\nRecognition (ASR) for a variety of reasons, including the lack of code-switched\ntraining data. Recently, we showed that monolingual ASR systems fine-tuned on\ncode-switched data deteriorate in performance on monolingual speech\nrecognition, which is not desirable as ASR systems deployed in multilingual\nscenarios should recognize both monolingual and code-switched speech with high\naccuracy. Our experiments indicated that this loss in performance could be\nmitigated by using certain strategies for fine-tuning and regularization,\nleading to improvements in both monolingual and code-switched ASR. In this\nwork, we present further improvements over our previous work by using domain\nadversarial learning to train task agnostic models. We evaluate the\nclassification accuracy of an adversarial discriminator and show that it can\nlearn shared layer parameters that are task agnostic. We train end-to-end ASR\nsystems starting with a pooled model that uses monolingual and code-switched\ndata along with the adversarial discriminator. Our proposed technique leads to\nreductions in Word Error Rates (WER) in monolingual and code-switched test sets\nacross three language pairs.\n