ESPnet: End-to-End Speech Processing Toolkit
2018/03/30 by Shinji Watanabe, Takaaki Hori, Watanabe, Shinji +21 · 53 citations
Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Music and Audio Processing #Natural Language Processing Techniques #Speech Recognition and Synthesis
paper · pdf · doi:10.48550/arxiv.1804.00015
openalex publication_date 2018/03/30 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Abstract
This paper introduces a new open source platform for end-to-end speech processing named ESPnet. ESPnet mainly focuses on end-to-end automatic speech recognition (ASR), and adopts widely-used dynamic neural network toolkits, Chainer and PyTorch, as a main deep learning engine. ESPnet also follows the Kaldi ASR toolkit style for data processing, feature extraction/format, and recipes to provide a complete setup for speech recognition and other speech processing experiments. This paper explains a major architecture of this software platform, several important functionalities, which differentiate ESPnet from other open source ASR toolkits, and experimental results with major ASR benchmarks.
Citations
Cited by
- Incorporating Error Level Noise Embedding for Improving LLM-Assisted Robustness in Persian Speech Recognition
- Hearing to Translate: The Effectiveness of Speech Modality Integration into LLMs
- Robust Training of Singing Voice Synthesis Using Prior and Posterior Uncertainty
- All-in-One ASR: Unifying Encoder-Decoder Models of CTC, Attention, and Transducer in Dual-Mode ASR
- Comparing Unsupervised and Supervised Semantic Speech Tokens: A Case Study of Child ASR
- VSpeechLM: A Visual Speech Language Model for Visual Text-to-Speech Task
- Building Robust and Scalable Multilingual ASR for Indian Languages
- TEDxTN: A Three-way Speech Translation Corpus for Code-Switched Tunisian Arabic - English
- Towards Effective and Efficient Non-autoregressive decoders for Conformer and LLM-based ASR using Block-based Attention Mask
- Towards General Auditory Intelligence: Large Multimodal Models for Machine Listening and Speaking
- BSCodec: A Band-Split Neural Codec for High-Quality Universal Audio Reconstruction
- Dissecting Sensitivity to Training Language in Self-Supervised Speech Learning Using Neural Audio Codec Tokens
- LSSED: a large-scale dataset and benchmark for speech emotion recognition
- Semantic-WER: A Unified Metric for the Evaluation of ASR Transcript for End Usability
- CIF: Continuous Integrate-and-Fire for End-to-End Speech Recognition
- POWSM: A Phonetic Open Whisper-Style Speech Foundation Model
- A Neural Model for Contextual Biasing Score Learning and Filtering
- ReFESS-QI: Reference-Free Evaluation For Speech Separation With Joint Quality And Intelligibility Scoring
- Language model integration based on memory control for sequence to sequence speech recognition
- Chain-of-Thought Reasoning in Streaming Full-Duplex End-to-End Spoken Dialogue Systems
- Automatic Speech Recognition in the Modern Era: Architectures, Training, and Evaluation
- Open ASR Leaderboard: Towards Reproducible and Transparent Multilingual and Long-Form Speech Recognition Evaluation
- TokenChain: A Discrete Speech Chain via Semantic Token Modeling
- Baseline Systems For The 2025 Low-Resource Audio Codec Challenge
- VSSFlow: Unifying Video-conditioned Sound and Speech Generation via Joint Learning
- Code-switching Speech Recognition Under the Lens: Model- and Data-Centric Perspectives
- Lightweight Front-end Enhancement for Robust ASR via Frame Resampling and Sub-Band Pruning
- Objective Evaluation of Prosody and Intelligibility in Speech Synthesis via Conditional Prediction of Discrete Tokens
- UMA-Split: unimodal aggregation for both English and Mandarin non-autoregressive speech recognition
- CS-FLEURS: A Massively Multilingual and Code-Switched Speech Dataset
- GLAD: Global-Local Aware Dynamic Mixture-of-Experts for Multi-Talker ASR
- WeNet: Production oriented Streaming and Non-streaming End-to-End Speech Recognition Toolkit
- The ASRU 2019 Mandarin-English Code-Switching Speech Recognition Challenge: Open Datasets, Tracks, Methods and Results
- Deep Discriminative Feature Learning for Accent Recognition
- Geolocation-Aware Robust Spoken Language Identification
- ESPnet2-TTS: Extending the Edge of TTS Research
- Unified Learnable 2D Convolutional Feature Extraction for ASR
- A Bottom-up Framework with Language-universal Speech Attribute Modeling for Syllable-based ASR
- Benchmarking Training Paradigms, Dataset Composition, and Model Scaling for Child ASR in ESPnet
- Graph Connectionist Temporal Classification for Phoneme Recognition
- SSVD: Structured SVD for Parameter-Efficient Fine-Tuning and Benchmarking under Domain Shift in ASR
- Towards Improved Speech Recognition through Optimized Synthetic Data Generation
- Benchmarking Large Pretrained Multilingual Models on Québec French Speech Recognition
- Unifying Diarization, Separation, and ASR with Multi-Speaker Encoder
- CAMÕES: A Comprehensive Automatic Speech Recognition Benchmark for European Portuguese
- DFANet: Deep Feature Aggregation for Real-Time Semantic Segmentation
- Exploring Cross-Utterance Speech Contexts for Conformer-Transducer Speech Recognition Systems
- Selection of Layers from Self-supervised Learning Models for Predicting Mean-Opinion-Score of Speech
- DeCRED: Decoder-Centric Regularization for Encoder-Decoder Based Speech Recognition
- Word Error Rate Definitions and Algorithms for Long-Form Multi-talker Speech Recognition
- Advancing Speech Quality Assessment Through Scientific Challenges and Open-source Activities
- Cross-Attention End-to-End ASR for Two-Party Conversations
- Dialog-context aware end-to-end speech recognition
Related