vix.ing · top · new · best · stats · spec

ESPnet: End-to-End Speech Processing Toolkit

2018/03/30 by Shinji Watanabe, Takaaki Hori, Watanabe, Shinji +21 · 53 citations
Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Music and Audio Processing #Natural Language Processing Techniques #Speech Recognition and Synthesis

paper · pdf · doi:10.48550/arxiv.1804.00015

openalex publication_date 2018/03/30 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28

Abstract

This paper introduces a new open source platform for end-to-end speech processing named ESPnet. ESPnet mainly focuses on end-to-end automatic speech recognition (ASR), and adopts widely-used dynamic neural network toolkits, Chainer and PyTorch, as a main deep learning engine. ESPnet also follows the Kaldi ASR toolkit style for data processing, feature extraction/format, and recipes to provide a complete setup for speech recognition and other speech processing experiments. This paper explains a major architecture of this software platform, several important functionalities, which differentiate ESPnet from other open source ASR toolkits, and experimental results with major ASR benchmarks.

Citations

Cited by

Related