vix.ing · top · new · best · stats

Multilingual Speech Recognition With A Single End-To-End Model

2017/11/06 by Shubham Toshniwal, Toshniwal, Shubham, Tara N. Sainath +12 · 8 citations
Computer Science · Engineering · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Natural Language Processing Techniques #Speech Recognition and Synthesis #cs.AI #cs.CL #eess.AS #electronic engineering #information engineering

paper · pdf · doi:10.48550/arxiv.1711.01694

Accepted in ICASSP 2018

openalex publication_date 2017/11/06 · arxiv created 2018/02/15 · arxiv updated 2018/02/16 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28

Abstract

Training a conventional automatic speech recognition (ASR) system to support multiple languages is challenging because the sub-word unit, lexicon and word inventories are typically language specific. In contrast, sequence-to-sequence models are well suited for multilingual ASR because they encapsulate an acoustic, pronunciation and language model jointly in a single network. In this work we present a single sequence-to-sequence ASR model trained on 9 different Indian languages, which have very little overlap in their scripts. Specifically, we take a union of language-specific grapheme sets and train a grapheme-based sequence-to-sequence model jointly on data from all languages. We find that this model, which is not explicitly given any information about language identity, improves recognition performance by 21% relative compared to analogous sequence-to-sequence models trained on each language individually. By modifying the model to accept a language identifier as an additional input feature, we further improve performance by an additional 7% relative and eliminate confusion between different languages.

Cited by

Related