vix.ing · top · new · best · stats

Deep Speech 2: End-to-End Speech Recognition in English and Mandarin

2015/12/08 by Dario Amodei, Rishita Anubhai, Amodei, Dario +66 · 1 voice · 2,182 citations
Computer Science · #Artificial intelligence #Computer network #Computer science #Deep learning #End-to-end principle #Key (lock) #Latency (audio) #Low latency (capital markets) #Mandarin Chinese #Natural Language Processing Techniques #Operating system #Parallel computing #Speech Recognition and Synthesis #Speech recognition #Speedup #Telecommunications #Topic Modeling #Variety (cybernetics) #cs.CL

paper · pdf · doi:10.48550/arxiv.1512.02595

published in arXiv (Cornell University) (Cornell University)

arxiv created 2015/12/08 · openalex publication_date 2015/12/08 · arxiv updated 2015/12/09 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28

Abstract

We show that an end-to-end deep learning approach can be used to recognize either English or Mandarin Chinese speech--two vastly different languages. Because it replaces entire pipelines of hand-engineered components with neural networks, end-to-end learning allows us to handle a diverse variety of speech including noisy environments, accents and different languages. Key to our approach is our application of HPC techniques, resulting in a 7x speedup over our previous system. Because of this efficiency, experiments that previously took weeks now run in days. This enables us to iterate more quickly to identify superior architectures and algorithms. As a result, in several cases, our system is competitive with the transcription of human workers when benchmarked on standard datasets. Finally, using a technique called Batch Dispatch with GPUs in the data center, we show that our system can be inexpensively deployed in an online setting, delivering low latency when serving users at scale.

Citations

Cited by

Discussions

Related