Deep Speech 2: End-to-End Speech Recognition in English and Mandarin
2015/12/08 by Dario Amodei, Rishita Anubhai, Amodei, Dario +66 · 1 voice · 2,182 citations
Computer Science · #Artificial intelligence #Computer network #Computer science #Deep learning #End-to-end principle #Key (lock) #Latency (audio) #Low latency (capital markets) #Mandarin Chinese #Natural Language Processing Techniques #Operating system #Parallel computing #Speech Recognition and Synthesis #Speech recognition #Speedup #Telecommunications #Topic Modeling #Variety (cybernetics) #cs.CL
paper · pdf · doi:10.48550/arxiv.1512.02595
published in arXiv (Cornell University) (Cornell University)
arxiv created 2015/12/08 · openalex publication_date 2015/12/08 · arxiv updated 2015/12/09 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Abstract
We show that an end-to-end deep learning approach can be used to recognize either English or Mandarin Chinese speech--two vastly different languages. Because it replaces entire pipelines of hand-engineered components with neural networks, end-to-end learning allows us to handle a diverse variety of speech including noisy environments, accents and different languages. Key to our approach is our application of HPC techniques, resulting in a 7x speedup over our previous system. Because of this efficiency, experiments that previously took weeks now run in days. This enables us to iterate more quickly to identify superior architectures and algorithms. As a result, in several cases, our system is competitive with the transcription of human workers when benchmarked on standard datasets. Finally, using a technique called Batch Dispatch with GPUs in the data center, we show that our system can be inexpensively deployed in an online setting, delivering low latency when serving users at scale.
Citations
Cited by
- From Profiling to Parameterization: Physics-Guided Acoustic Eavesdropping via Smartphone Accelerometers
- Event Extraction in Large Language Model
- Asynchronous Pipeline Parallelism for Real-Time Multilingual Lip Synchronization in Video Communication Systems
- MADTempo: An Interactive System for Multi-Event Temporal Video Retrieval with Query Augmentation
- Better audio representations are more brain-like: linking model-brain alignment with performance in downstream auditory tasks
- Omni-AVSR: Towards Unified Multimodal Speech Recognition with Large Language Models
- audio2chart: End to End Audio Transcription into playable Guitar Hero charts
- ATLAS: Adaptive Transfer Scaling Laws for Multilingual Pretraining, Finetuning, and Decoding the Curse of Multilinguality
- Policy Transfer for Continuous-Time Reinforcement Learning: A (Rough) Differential Equation Approach
- Proprioceptive Image: An Image Representation of Proprioceptive Data from Quadruped Robots for Contact Estimation Learning
- Multi-Agent Design Assistant for the Simulation of Inertial Fusion Energy
- An Explorative Study on Distributed Computing Techniques in Training and Inference of Large Language Models
- Spiralformer: Low Latency Encoder for Streaming Speech Recognition with Circular Layer Skipping and Early Exiting
- An overview of neural architectures for self-supervised audio representation learning from masked spectrograms
- Dynamic Spectrum Matching with One-shot Learning
- Recurrent Neural Networks With Limited Numerical Precision
- Memory Visualization for Gated Recurrent Neural Networks in Speech Recognition
- WaveGuard: Understanding and Mitigating Audio Adversarial Examples
- Experiments with Rich Regime Training for Deep Learning
- Can deep learning beat numerical weather prediction?
- Layer-wise Analysis for Quality of Multilingual Synthesized Speech
- Fully Distributed Multi-Robot Collision Avoidance via Deep Reinforcement Learning for Safe and Efficient Navigation in Complex Scenarios
- mixup: Beyond Empirical Risk Minimization
- Guided Source Separation Meets a Strong ASR Backend: Hitachi/Paderborn University Joint Investigation for Dinner Party ASR
- A Novel Fusion of Attention and Sequence to Sequence Autoencoders to Predict Sleepiness From Speech
- Towards Inclusive Communication: A Unified Framework for Generating Spoken Language from Sign, Lip, and Audio
- Improved Recurrent Neural Networks for Session-based Recommendations
- RNN-T Models Fail to Generalize to Out-of-Domain Audio: Causes and Solutions
- Do Explanations Reflect Decisions? A Machine-centric Strategy to Quantify the Performance of Explainability Algorithms
- Semantic Communications for Speech Recognition
- A Study of BFLOAT16 for Deep Learning Training
- CAAD 2018: Iterative Ensemble Adversarial Attack
- A Tutorial on Ultra-Reliable and Low-Latency Communications in 6G: Integrating Domain Knowledge into Deep Learning
- Advances in Joint CTC-Attention based End-to-End Speech Recognition with a Deep CNN Encoder and RNN-LM
- Transformer with Bidirectional Decoder for Speech Recognition
- Auxiliary Interference Speaker Loss for Target-Speaker Speech Recognition
- Two-stage Training for Chinese Dialect Recognition
- Continuous Transition: Improving Sample Efficiency for Continuous Control Problems via MixUp
- From Semi-supervised to Almost-unsupervised Speech Recognition with Very-low Resource by Jointly Learning Phonetic Structures from Audio and Text Embeddings
- WaveTTS: Tacotron-based TTS with Joint Time-Frequency Domain Loss
- Speech Recognition With No Speech Or With Noisy Speech Beyond English
- EmoAugNet: A Signal-Augmented Hybrid CNN-LSTM Framework for Speech Emotion Recognition
- Speech-to-LaTeX: New Models and Datasets for Converting Spoken Equations and Sentences
- 4K-Memristor Analog-Grade Passive Crossbar Circuit
- Multilingual Graphemic Hybrid ASR with Massive Data Augmentation
- RCR-AF: Enhancing Model Generalization via Rademacher Complexity Reduction Activation Function
- Improving Adversarial Robustness Through Adaptive Learning-Driven Multi-Teacher Knowledge Distillation
- Learning to See Inside Opaque Liquid Containers using Speckle Vibrometry
- Low Latency End-to-End Streaming Speech Recognition with a Scout Network
- A Simplified Fully Quantized Transformer for End-to-end Speech Recognition
- Machine Learning With Neuromorphic Photonics
- Transformer-based Online CTC/attention End-to-End Speech Recognition Architecture
- Quadratic Suffices for Over-parametrization via Matrix Chernoff Bound
- SGAD: Soft-Guided Adaptively-Dropped Neural Network
- Delta Networks for Optimized Recurrent Network Computation
- Unsupervised Speech Recognition
- Cycle-consistency training for end-to-end speech recognition
- Label-Synchronous Speech-to-Text Alignment for ASR Using Forward and Backward Transformers
- Improving RNN Transducer Modeling for End-to-End Speech Recognition
- Online Evolutionary Batch Size Orchestration for Scheduling Deep Learning Workloads in GPU Clusters
- ReconVAT: A Semi-Supervised Automatic Music Transcription Framework for Low-Resource Real-World Data
- The History of Speech Recognition to the Year 2030
- A Spectral Energy Distance for Parallel Speech Synthesis
- Theory of the Frequency Principle for General Deep Neural Networks
- Hierarchical Multitask Learning for CTC-based Speech Recognition
- Residual Convolutional CTC Networks for Automatic Speech Recognition
- DARTS: Dialectal Arabic Transcription System
- Insertion-Based Modeling for End-to-End Automatic Speech Recognition
- Speech-Based Visual Question Answering
- Deep Triphone Embedding Improves Phoneme Recognition
- Boosting Active Learning for Speech Recognition with Noisy Pseudo-labeled Samples
- On Large Batch Training and Sharp Minima: A Fokker-Planck Perspective
- Unified Streaming and Non-streaming Two-pass End-to-end Model for Speech Recognition
- Application of Word2vec in Phoneme Recognition
- Exploiting Contextual Information with Deep Neural Networks
- An investigation of phone-based subword units for end-to-end speech recognition
- An improved hybrid CTC-Attention model for speech recognition
- Robust Watermarking of Neural Network with Exponential Weighting
- Universal adversarial examples in speech command classification
- MIMO-SPEECH: End-to-End Multi-Channel Multi-Speaker Speech Recognition
- SimulLR: Simultaneous Lip Reading Transducer with Attention-Guided Adaptive Memory
- Active Learning for Speech Recognition: the Power of Gradients
- Task Agnostic Continual Learning Using Online Variational Bayes
- Explaining the Attention Mechanism of End-to-End Speech Recognition Using Decision Trees
- M2DAO-Talker: Harmonizing Multi-granular Motion Decoupling and Alternating Optimization for Talking-head Generation
- Advancing CTC-CRF Based End-to-End Speech Recognition with Wordpieces and Conformers
- Omni-Router: Sharing Routing Decisions in Sparse Mixture-of-Experts for Speech Recognition
- Adaptive Dense-to-Sparse Paradigm for Pruning Online Recommendation System with Non-Stationary Data
- Improved Regularization Techniques for End-to-End Speech Recognition
- Stochastic Mirror Descent on Overparameterized Nonlinear Models: Convergence, Implicit Regularization, and Generalization
- Understanding Modern Techniques in Optimization: Frank-Wolfe, Nesterov's Momentum, and Polyak's Momentum
- Towards End-to-end Automatic Code-Switching Speech Recognition
- Multi-Modal Emotion recognition on IEMOCAP Dataset using Deep Learning
- Speeding up Deep Model Training by Sharing Weights and Then Unsharing
- Dual Head Adversarial Training
- Unbounded cache model for online language modeling with open vocabulary
- A Brief Survey and an Application of Semantic Image Segmentation for Autonomous Driving
- i-Mix: A Domain-Agnostic Strategy for Contrastive Representation Learning
- Dynamic Encoder Transducer: A Flexible Solution For Trading Off Accuracy For Latency
- Letter-Based Speech Recognition with Gated ConvNets
- Extending Recurrent Neural Aligner for Streaming End-to-End Speech Recognition in Mandarin
- XCloud: Design and Implementation of AI Cloud Platform with RESTful API Service
- A Microprocessor implemented in 65nm CMOS with Configurable and Bit-scalable Accelerator for Programmable In-memory Computing
- A SOT-MRAM-based Processing-In-Memory Engine for Highly Compressed DNN Implementation
- Analyzing and Fine-Tuning Whisper Models for Multilingual Pilot Speech Transcription in the Cockpit
- The Architectural Implications of Facebook's DNN-based Personalized Recommendation
- Sequence Modeling via Segmentations
- Removing Backdoor-Based Watermarks in Neural Networks with Limited Data
- MEL: Multi-level Ensemble Learning for Resource-Constrained Environments
- Deceiving End-to-End Deep Learning Malware Detectors using Adversarial Examples
- Compute Trends Across Three Eras of Machine Learning
- Houdini: Fooling Deep Structured Prediction Models
- Toward Streaming ASR with Non-Autoregressive Insertion-based Model
- Benchmarking TPU, GPU, and CPU Platforms for Deep Learning
- Early Attentive Sparsification Accelerates Neural Speech Transcription
- An Online Attention-based Model for Speech Recognition
- SyncTalk++: High-Fidelity and Efficient Synchronized Talking Heads Synthesis Using Gaussian Splatting
- Resource aware design of a deep convolutional-recurrent neural network for speech recognition through audio-visual sensor fusion
- Exploring Neural Transducers for End-to-End Speech Recognition
- Improving End-to-End Speech Recognition with Policy Learning
- Formant Tracking Using Dilated Convolutional Networks Through Dense Connection with Gating Mechanism
- Towards End-to-End Earthquake Monitoring Using a Multitask Deep Learning Model
- Phonemic and Graphemic Multilingual CTC Based Speech Recognition
- Learning from Learning Machines: Optimisation, Rules, and Social Norms
- Machine Learning for Robust Identification of Complex Nonlinear Dynamical Systems: Applications to Earth Systems Modeling
- FPGA-Based Low-Power Speech Recognition with Recurrent Neural Networks
- A comparable study of modeling units for end-to-end Mandarin speech recognition
- Attention-Augmented End-to-End Multi-Task Learning for Emotion Prediction from Speech
- Masked Pre-trained Encoder base on Joint CTC-Transformer
- A Deep-learning-based Method for PIR-based Multi-person Localization
- Faster, Simpler and More Accurate Hybrid ASR Systems Using Wordpieces
- Fully Convolutional Speech Recognition
- Jasper: An End-to-End Convolutional Neural Acoustic Model
- End-to-end ASR: from Supervised to Semi-Supervised Learning with Modern Architectures
- A systematic comparison of grapheme-based vs. phoneme-based label units for encoder-decoder-attention models
- PyChain: A Fully Parallelized PyTorch Implementation of LF-MMI for End-to-End ASR
- Improved training for online end-to-end speech recognition systems
- Learning Robust and Multilingual Speech Representations
- Attention-based ASR with Lightweight and Dynamic Convolutions
- Advancing Multi-Accented LSTM-CTC Speech Recognition using a Domain Specific Student-Teacher Learning Paradigm
- Refining Automatic Speech Recognition System for older adults
- End-to-end named entity extraction from speech
- power-law nonlinearity with maximally uniform distribution criterion for improved neural network training in automatic speech recognition
- Research on Modeling Units of Transformer Transducer for Mandarin Speech Recognition
- Joint Regularization on Activations and Weights for Efficient Neural Network Pruning
- Towards thinner convolutional neural networks through Gradually Global Pruning
- Mixture of Expert/Imitator Networks: Scalable Semi-supervised Learning Framework
- Efficient Neural and Numerical Methods for High-Quality Online Speech Spectrogram Inversion via Gradient Theorem
- Accurate and Energy-Efficient Classification with Spiking Random Neural Network: Corrected and Expanded Version
- The Case for Strong Scaling in Deep Learning: Training Large 3D CNNs with Hybrid Parallelism
- Deep Learning at 15PF: Supervised and Semi-Supervised Classification for Scientific Data
- A novel pyramidal-FSMN architecture with lattice-free MMI for speech recognition
- No-regret Non-convex Online Meta-Learning
- SF-Net: Structured Feature Network for Continuous Sign Language Recognition
- Latency-Controlled Neural Architecture Search for Streaming Speech Recognition
- Learning Recurrent Binary/Ternary Weights
- SEC4SR: A Security Analysis Platform for Speaker Recognition
- Dub-S2ST: Textless Speech-to-Speech Translation for Seamless Dubbing
- Towards Online End-to-end Transformer Automatic Speech Recognition
- Two-stage Textual Knowledge Distillation for End-to-End Spoken Language Understanding
- End-to-End Bengali Speech Recognition
- TabSTAR: A Tabular Foundation Model for Tabular Data with Text Fields
- Hard Sample Mining for the Improved Retraining of Automatic Speech Recognition
- A Simple yet Effective Baseline for Robust Deep Learning with Noisy Labels
- Sequence-Level Knowledge Distillation for Model Compression of Attention-based Sequence-to-Sequence Speech Recognition
- MaskCycleGAN-VC: Learning Non-parallel Voice Conversion with Filling in Frames
- Redox: Improving I/O Efficiency of Model Training Through File Redirection
- Snore-GANs: Improving Automatic Snore Sound Classification with Synthesized Data
- Rethinking Full Connectivity in Recurrent Neural Networks
- Straggler-Resilient Distributed Machine Learning with Dynamic Backup Workers
- Scaling and Enhancing LLM-based AVSR: A Sparse Mixture of Projectors Approach
- Hard Class Rectification for Domain Adaptation
- TBD: Benchmarking and Analyzing Deep Neural Network Training
- Towards Learning to Speak and Hear Through Multi-Agent Communication\n over a Continuous Acoustic Channel
- Wav2Letter: an End-to-End ConvNet-based Speech Recognition System
- Improving LSTM-CTC based ASR performance in domains with limited training data
- Improving speech recognition models with small samples for air traffic control systems
- Guiding CTC Posterior Spike Timings for Improved Posterior Fusion and Knowledge Distillation
- Fooling OCR Systems with Adversarial Text Images
- Reinforcement Learning and Adaptive Sampling for Optimized DNN Compilation
- AIBench: An Industry Standard Internet Service AI Benchmark Suite
- Multi-task Recurrent Model for True Multilingual Speech Recognition
- Adversarial Regression with Doubly Non-negative Weighting Matrices
- Edge AIBench: Towards Comprehensive End-to-end Edge Computing Benchmarking
- LipNet: End-to-End Sentence-level Lipreading
- Halo: Learning Semantics-Aware Representations for Cross-Lingual Information Extraction
- Multi-Head Decoder for End-to-End Speech Recognition
- Self-Adaptive Reconfigurable Arrays (SARA): Using ML to Assist Scaling GEMM Acceleration
- Deep Recurrent Convolutional Neural Network: Improving Performance For Speech Recognition
- Training Deep Neural Networks Using Posit Number System
- Efficient conformer-based speech recognition with linear attention
- Gated Recurrent Unit Based Acoustic Modeling with Future Context
- Bayesian Sparsification of Recurrent Neural Networks
- Demystifying Parallel and Distributed Deep Learning
- Graphcore C2 Card performance for image-based deep learning application: A Report
- A Dynamically Controlled Recurrent Neural Network for Modeling Dynamical Systems
- Knowledge Distillation from BERT Transformer to Speech Transformer for Intent Classification
- Deep segmental phonetic posterior-grams based discovery of non-categories in L2 English speech
- Exponential Moving Average Model in Parallel Speech Recognition Training
- Modeling the Second Player in Distributionally Robust Optimization
- Landscape and training regimes in deep learning
- RT-RCG: Neural Network and Accelerator Search Towards Effective and Real-time ECG Reconstruction from Intracardiac Electrograms
- AIBench: An Agile Domain-specific Benchmarking Methodology and an AI Benchmark Suite
- Variational hybridization and transformation for large inaccurate noisy-or networks
- Training Neural Speech Recognition Systems with Synthetic Speech Augmentation
- SparseNN: An Energy-Efficient Neural Network Accelerator Exploiting Input and Output Sparsity
- End-to-End Deep Fault-Tolerant Control
- Benchmarking the Performance and Energy Efficiency of AI Accelerators for AI Training
- Learning spectro-temporal representations of complex sounds with parameterized neural networks
- Speech recognition for air traffic control via feature learning and end-to-end training
- Effects of Number of Filters of Convolutional Layers on Speech Recognition Model Accuracy
- Human-Machine Collaborative Design for Accelerated Design of Compact Deep Neural Networks for Autonomous Driving
- ARM-Net: Adaptive Relation Modeling Network for Structured Data
- Did you hear that? Adversarial Examples Against Automatic Speech Recognition
- Improving Noise Robustness of an End-to-End Neural Model for Automatic Speech Recognition
- End-To-End Speech Recognition Using A High Rank LSTM-CTC Based Model
- Attention-Based End-to-End Speech Recognition on Voice Search
- Insights on Neural Representations for End-to-End Speech Recognition
- Human-Machine Interaction Speech Corpus from the ROBIN project
- PDAugment: Data Augmentation by Pitch and Duration Adjustments for Automatic Lyrics Transcription
- ClovaCall: Korean Goal-Oriented Dialog Speech Corpus for Automatic Speech Recognition of Contact Centers
- Towards Automatic Face-to-Face Translation
- Learning Without Feedback: Fixed Random Learning Signals Allow for Feedforward Training of Deep Neural Networks
- Learning Shared Encoding Representation for End-to-End Speech\n Recognition Models
- Direct Speech-to-Image Translation
- Recurrent Batch Normalization
- Distributed Deep Learning Strategies For Automatic Speech Recognition
- GaDei: On Scale-up Training As A Service For Deep Learning
- CNN-based MultiChannel End-to-End Speech Recognition for everyday home environments
- Poem Meter Classification of Recited Arabic Poetry: Integrating High-Resource Systems for a Low-Resource Task
- Beam Search Decoding using Manner of Articulation Detection Knowledge Derived from Connectionist Temporal Classification
- xDeepFM
- Speech recognition [wikipedia]
Discussions
Related