2019/07/15 by Kunal Dhawan, Dhawan, Kunal, Ganji Sreeram +5 · 1 citation
Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
paper · pdf · doi:10.48550/arxiv.1907.08293
openalex publication_date 2019/07/15 · openalex created_date 2021/05/24 · openalex updated_date 2026/07/28
End-to-end (E2E) systems are fast replacing the conventional systems in the\ndomain of automatic speech recognition. As the target labels are learned\ndirectly from speech data, the E2E systems need a bigger corpus for effective\ntraining. In the context of code-switching task, the E2E systems face two\nchallenges: (i) the expansion of the target set due to multiple languages\ninvolved, and (ii) the lack of availability of sufficiently large\ndomain-specific corpus. Towards addressing those challenges, we propose an\napproach for reducing the number of target labels for reliable training of the\nE2E systems on limited data. The efficacy of the proposed approach has been\ndemonstrated on two prominent architectures, namely CTC-based and\nattention-based E2E networks. The experimental validations are performed on a\nrecently created Hindi-English code-switching corpus. For contrast purpose, the\nresults for the full target set based E2E system and a hybrid DNN-HMM system\nare also reported.\n