2020/11/02 by Ching-Feng Yeh, Yongqiang Wang, Yeh, Ching-Feng +11
Computer Science · #Advanced Neural Network Applications #Computation and Language (cs.CL) #FOS: Computer and information sciences #Natural Language Processing Techniques #Speech Recognition and Synthesis
paper · pdf · doi:10.48550/arxiv.2011.07120
openalex publication_date 2020/11/02 · openalex created_date 2022/07/25 · openalex updated_date 2026/07/28
Attention-based models have been gaining popularity recently for their strong\nperformance demonstrated in fields such as machine translation and automatic\nspeech recognition. One major challenge of attention-based models is the need\nof access to the full sequence and the quadratically growing computational cost\nconcerning the sequence length. These characteristics pose challenges,\nespecially for low-latency scenarios, where the system is often required to be\nstreaming. In this paper, we build a compact and streaming speech recognition\nsystem on top of the end-to-end neural transducer architecture with\nattention-based modules augmented with convolution. The proposed system equips\nthe end-to-end models with the streaming capability and reduces the large\nfootprint from the streaming attention-based model using augmented memory. On\nthe LibriSpeech dataset, our proposed system achieves word error rates 2.7% on\ntest-clean and 5.8% on test-other, to our best knowledge the lowest among\nstreaming approaches reported so far.\n