2019/03/28 by Shane Settle, Settle, Shane, Kartik Audhkhasi +5
Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Music and Audio Processing #Speech Recognition and Synthesis #Speech and Audio Processing #Topic Modeling
paper · pdf · doi:10.48550/arxiv.1903.12306
openalex publication_date 2019/03/28 · openalex created_date 2022/07/24 · openalex updated_date 2026/07/28
Direct acoustics-to-word (A2W) systems for end-to-end automatic speech\nrecognition are simpler to train, and more efficient to decode with, than\nsub-word systems. However, A2W systems can have difficulties at training time\nwhen data is limited, and at decoding time when recognizing words outside the\ntraining vocabulary. To address these shortcomings, we investigate the use of\nrecently proposed acoustic and acoustically grounded word embedding techniques\nin A2W systems. The idea is based on treating the final pre-softmax weight\nmatrix of an AWE recognizer as a matrix of word embedding vectors, and using an\nexternally trained set of word embeddings to improve the quality of this\nmatrix. In particular we introduce two ideas: (1) Enforcing similarity at\ntraining time between the external embeddings and the recognizer weights, and\n(2) using the word embeddings at test time for predicting out-of-vocabulary\nwords. Our word embedding model is acoustically grounded, that is it is learned\njointly with acoustic embeddings so as to encode the words' acoustic-phonetic\ncontent; and it is parametric, so that it can embed any arbitrary (potentially\nout-of-vocabulary) sequence of characters. We find that both techniques improve\nthe performance of an A2W recognizer on conversational telephone speech.\n