vix.ing · top · new · best · stats · spec

Direct Acoustics-to-Word Models for English Conversational Speech\n Recognition

2017/03/22 by Kartik Audhkhasi, Bhuvana Ramabhadran, Audhkhasi, Kartik +7 · 2 citations
Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (stat.ML) #Music and Audio Processing #Natural Language Processing Techniques #Neural and Evolutionary Computing (cs.NE) #Speech Recognition and Synthesis

paper · pdf · doi:10.48550/arxiv.1703.07754

openalex publication_date 2017/03/22 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28

Abstract

Recent work on end-to-end automatic speech recognition (ASR) has shown that\nthe connectionist temporal classification (CTC) loss can be used to convert\nacoustics to phone or character sequences. Such systems are used with a\ndictionary and separately-trained Language Model (LM) to produce word\nsequences. However, they are not truly end-to-end in the sense of mapping\nacoustics directly to words without an intermediate phone representation. In\nthis paper, we present the first results employing direct acoustics-to-word CTC\nmodels on two well-known public benchmark tasks: Switchboard and CallHome.\nThese models do not require an LM or even a decoder at run-time and hence\nrecognize speech with minimal complexity. However, due to the large number of\nword output units, CTC word models require orders of magnitude more data to\ntrain reliably compared to traditional systems. We present some techniques to\nmitigate this issue. Our CTC word model achieves a word error rate of\n13.0%/18.8% on the Hub5-2000 Switchboard/CallHome test sets without any LM or\ndecoder compared with 9.6%/16.0% for phone-based CTC with a 4-gram LM. We also\npresent rescoring results on CTC word model lattices to quantify the\nperformance benefits of a LM, and contrast the performance of word and phone\nCTC models.\n

Citations

Cited by

Related