2018/10/28 by Hila Gonen, Yoav Goldberg, Gonen, Hila +1
Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Natural Language Processing Techniques #Speech Recognition and Synthesis #Speech and dialogue systems
paper · pdf · doi:10.48550/arxiv.1810.11895
openalex publication_date 2018/10/28 · openalex created_date 2022/03/02 · openalex updated_date 2026/07/28
We focus on the problem of language modeling for code-switched language, in\nthe context of automatic speech recognition (ASR). Language modeling for\ncode-switched language is challenging for (at least) three reasons: (1) lack of\navailable large-scale code-switched data for training; (2) lack of a replicable\nevaluation setup that is ASR directed yet isolates language modeling\nperformance from the other intricacies of the ASR system; and (3) the reliance\non generative modeling. We tackle these three issues: we propose an\nASR-motivated evaluation setup which is decoupled from an ASR system and the\nchoice of vocabulary, and provide an evaluation dataset for English-Spanish\ncode-switching. This setup lends itself to a discriminative training approach,\nwhich we demonstrate to work better than generative language modeling. Finally,\nwe explore a variety of training protocols and verify the effectiveness of\ntraining with large amounts of monolingual data followed by fine-tuning with\nsmall amounts of code-switched data, for both the generative and discriminative\ncases.\n