Kaldi Speech Recognition Toolkit
2024/01/01 by Daniel Povey
Computer Science · #Music and Audio Processing #Speech Recognition and Synthesis #Speech and Audio Processing
paper · pdf · doi:10.57702/jb3fvbn9
openalex created_date 2016/06/24 · openalex publication_date 2024/01/01 · openalex updated_date 2026/07/23
Abstract
Abstract—We describe the design of Kaldi, a free, open-source toolkit for speech recognition research. Kaldi provides a speech recognition system based on finite-state transducers (using the freely available OpenFst), together with detailed documentation and scripts for building complete recognition systems. Kaldi is written is C++, and the core library supports modeling of arbitrary phonetic-context sizes, acoustic modeling with subspace Gaussian mixture models (SGMM) as well as standard Gaussian mixture models, together with all commonly used linear and affine transforms. Kaldi is released under the Apache License v2.0, which is highly nonrestrictive, making it suitable for a wide community of users. I.
Cited by
- ASVspoof 2019: A large-scale public database of synthesized, converted and replayed speech
- The third ‘CHiME’ speech separation and recognition challenge: Analysis and outcomes
- Characterisation of voice quality of Parkinson’s disease using differential phonological posterior features
- Knowledge Distillation For Recurrent Neural Network Language Modeling With Trust Regularization
- ESPnet: End-to-End Speech Processing Toolkit
- Sequential Routing Framework: Fully Capsule Network-based Speech Recognition
- A Survey on Deep Learning Toolkits and Libraries for Intelligent User Interfaces
- CHiME-6 Challenge:Tackling Multispeaker Speech Recognition for Unsegmented Recordings
- Adapting GPT, GPT-2 and BERT Language Models for Speech Recognition
- ESE: Efficient Speech Recognition Engine with Sparse LSTM on FPGA
- Analysis of Length Normalization in End-to-End Speaker Verification System
- Fast Transformers with Clustered Attention
- End-to-End Neural Speaker Diarization with Self-attention
- Masks Fusion with Multi-Target Learning For Speech Enhancement
- Mandarin tone modeling using recurrent neural networks
- Deep Learning in Robotics: A Review of Recent Research
- AV Speech Enhancement Challenge using a Real Noisy Corpus
- Design and Optimization of a Speech Recognition Front-End for Distant-Talking Control of a Music Playback Device
- Augmenting Librispeech with French Translations: A Multimodal Corpus for Direct Speech Translation Evaluation
- Spoken Language Translation for Polish
- Emformer: Efficient Memory Transformer Based Acoustic Model For Low Latency Streaming Speech Recognition
- Detection of Consonant Errors in Disordered Speech Based on Consonant-vowel Segment Embedding
- Truncated Variational Expectation Maximization
- Essence Knowledge Distillation for Speech Recognition
- Measuring the Effectiveness of Voice Conversion on Speaker Identification and Automatic Speech Recognition Systems
- Investigation of Speaker-adaptation methods in Transformer based ASR
- Semantic-WER: A Unified Metric for the Evaluation of ASR Transcript for End Usability
- A Probabilistic Framework for Lexicon-based Keyword Spotting in Handwritten Text Images
- Accented Speech Recognition Inspired by Human Perception
- CIF: Continuous Integrate-and-Fire for End-to-End Speech Recognition
- Domain-Invariant Speaker Vector Projection by Model-Agnostic Meta-Learning
- Large-Scale Approximate Kernel Canonical Correlation Analysis
- Speaker Separation Using Speaker Inventories and Estimated Speech
- Advances in subword-based HMM-DNN speech recognition across languages
- The 2020 ESPnet update: new features, broadened applications, performance improvements, and future plans
- Semi-Supervised Model Training for Unbounded Conversational Speech Recognition
- American Sign Language fingerspelling recognition from video: Methods for unrestricted recognition and signer-independence
- Interpretable Filter Learning Using Soft Self-attention For Raw Waveform Speech Recognition
- Teach an all-rounder with experts in different domains
- Multi-channel Speech Enhancement with 2-D Convolutional Time-frequency Domain Features and a Pre-trained Acoustic Model
- Multistream CNN for Robust Acoustic Modeling
- Unified Signal Compression Using Generative Adversarial Networks
- DELTA: A DEep learning based Language Technology plAtform
- Self-and-Mixed Attention Decoder with Deep Acoustic Structure for Transformer-based LVCSR
- AP20-OLR Challenge: Three Tasks and Their Baselines
- Recognize Foreign Low-Frequency Words with Similar Pairs
- Reducing Exposure Bias in Training Recurrent Neural Network Transducers
- Multi-Graph Decoding for Code-Switching ASR
- Acoustic Feature Learning via Deep Variational Canonical Correlation Analysis
- Improving pronunciation assessment via ordinal regression with anchored reference samples
- The RWTH ASR System for TED-LIUM Release 2: Improving Hybrid HMM with SpecAugment
- Towards Relevance and Sequence Modeling in Language Recognition
- Kaleidoscope: An Efficient, Learnable Representation For All Structured Linear Maps
- Reward Augmented Maximum Likelihood for Neural Structured Prediction
- Unsupervised Learning of Disentangled and Interpretable Representations from Sequential Data
- THUEE system description for NIST 2019 SRE CTS Challenge
- Phonetic Temporal Neural Model for Language Identification
- Neural Polysynthetic Language Modelling
- Prosody Transfer in Neural Text to Speech Using Global Pitch and Loudness Features
- Partial AUC optimization based deep speaker embeddings with class-center learning for text-independent speaker verification
- Phone-aware Neural Language Identification
- Unsupervised Speaker Adaptation using Attention-based Speaker Memory for\n End-to-End ASR
- Deep Spiking Neural Networks for Large Vocabulary Automatic Speech Recognition
- Improving Cross-Lingual Transfer Learning for End-to-End Speech Recognition with Speech Translation
- Improved Meta-Learning Training for Speaker Verification
- Lightweight Speech Enhancement in Unseen Noisy and Reverberant Conditions using KISS-GEV Beamforming
- Exploring Gaussian mixture model framework for speaker adaptation of deep neural network acoustic models
- Investigating the role of L1 in automatic pronunciation evaluation of L2 speech
- An Objective Evaluation Framework for Pathological Speech Synthesis
- Optimising The Input Window Alignment in CD-DNN Based Phoneme Recognition for Low Latency Processing
- Combining Spatial Clustering with LSTM Speech Models for Multichannel\n Speech Enhancement
- VAE-based Domain Adaptation for Speaker Verification
- Recurrent Neural Networks With Limited Numerical Precision
- Memory Visualization for Gated Recurrent Neural Networks in Speech Recognition
- Acoustic data-driven lexicon learning based on a greedy pronunciation selection framework
- UIAI System for Short-Duration Speaker Verification Challenge 2020
- Effects of Language Relatedness for Cross-lingual Transfer Learning in Character-Based Language Models
- Multilingual End-to-End Speech Translation
- The ASRU 2019 Mandarin-English Code-Switching Speech Recognition Challenge: Open Datasets, Tracks, Methods and Results
- Time Domain Audio Visual Speech Separation
- PI-Edge: A Low-Power Edge Computing System for Real-Time Autonomous Driving Services
- Improving Structured Text Recognition with Regular Expression Biasing
- SpeechBrain: A General-Purpose Speech Toolkit
- SMS-WSJ: Database, performance measures, and baseline recipe for multi-channel source separation and recognition
- ESPnet2-TTS: Extending the Edge of TTS Research
- Adapting End-to-End Speech Recognition for Readable Subtitles
- Adversarial Attacks and Defenses for Speech Recognition Systems
- Real-time, Universal, and Robust Adversarial Attacks Against Speaker Recognition Systems
- Listen, Attend and Spell
- Asteroid: the PyTorch-based audio source separation toolkit for\n researchers
- On combining features for single-channel robust speech recognition in reverberant environments
- A language score based output selection method for multilingual speech\n recognition
- Transformer Language Models with LSTM-based Cross-utterance Information Representation
- Modular End-to-end Automatic Speech Recognition Framework for Acoustic-to-word Model
- An evaluation of word-level confidence estimation for end-to-end automatic speech recognition
- BUT System Description to VoxCeleb Speaker Recognition Challenge 2019
- Putting An End to End-to-End: Gradient-Isolated Learning of Representations
- Fluent Translations from Disfluent Speech in End-to-End Speech\n Translation
- Automatic evaluation of reading aloud performance in children
- DeepThin: A Self-Compressing Library for Deep Neural Networks
- Characterizing Types of Convolution in Deep Convolutional Recurrent Neural Networks for Robust Speech Emotion Recognition
- LSTM-TDNN with convolutional front-end for Dialect Identification in the 2019 Multi-Genre Broadcast Challenge
- Differentiable Weighted Finite-State Transducers
- General-Purpose Speech Representation Learning through a Self-Supervised Multi-Granularity Framework
- Building Bilingual and Code-Switched Voice Conversion with Limited Training Data Using Embedding Consistency Loss
- Text-Free Prosody-Aware Generative Spoken Language Modeling
- ASR error management for improving spoken language understanding
- Prosodic Features from Large Corpora of Child-Directed Speech as Predictors of the Age of Acquisition of Words
- Improved MVDR Beamforming Using LSTM Speech Models to Clean Spatial Clustering Masks
- A Comparison and Combination of Unsupervised Blind Source Separation Techniques
- Speaking Speed Control of End-to-End Speech Synthesis using Sentence-Level Conditioning
- Advances in Joint CTC-Attention based End-to-End Speech Recognition with a Deep CNN Encoder and RNN-LM
- Segmental Recurrent Neural Networks for End-to-end Speech Recognition
- Auxiliary Interference Speaker Loss for Target-Speaker Speech Recognition
- Far-Field Automatic Speech Recognition
- Processing Phoneme Specific Segments for Cleft Lip and Palate Speech Enhancement
- From Semi-supervised to Almost-unsupervised Speech Recognition with Very-low Resource by Jointly Learning Phonetic Structures from Audio and Text Embeddings
- An Investigation of Enhancing CTC Model for Triggered Attention-based Streaming ASR
- Machine Speech Chain with One-shot Speaker Adaptation
- Exploration of End-to-end Synthesisers forZero Resource Speech Challenge\n 2020
- Investigation of End-To-End Speaker-Attributed ASR for Continuous Multi-Talker Recordings
- Multilingual Graphemic Hybrid ASR with Massive Data Augmentation
- Towards Unsupervised Automatic Speech Recognition Trained by Unaligned Speech and Text only
- Multi-Modal Transformers Utterance-Level Code-Switching Detection
- Techniques for Vocabulary Expansion in Hybrid Speech Recognition Systems
- ESPnet-ST IWSLT 2021 Offline Speech Translation System
- The DKU-Duke-Lenovo System Description for the Third DIHARD Speech Diarization Challenge
- Streaming End-to-End ASR based on Blockwise Non-Autoregressive Models
- Improving End-To-End Modeling for Mispronunciation Detection with Effective Augmentation Mechanisms
- Multitask Learning with CTC and Segmental CRF for Speech Recognition
Related