Povey, Daniel
- MUSAN: A Music, Speech, and Noise Corpus
2015/10/28 by Snyder, David, Chen, Guoguo, Povey, Daniel · 77 citations
#FOS: Computer and information sciences #Sound (cs.SD)
- CHiME-6 Challenge:Tackling Multispeaker Speech Recognition for Unsegmented Recordings
2020/04/20 by Shinji Watanabe, Watanabe, Shinji, Michael Mandel +39 · 27 citations
Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- Libriheavy: a 50,000 hours ASR corpus with punctuation casing and context
2023/09/15 by Wei Kang, Kang, Wei, Xiaoyu Yang +12 · 35 citations
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- Zipformer: A faster and better encoder for automatic speech recognition
2023/10/17 by Zengwei Yao, Liyong Guo, Yao, Zengwei +14 · 24 citations
Computer Science · #Speech Recognition and Synthesis #Music and Audio Processing #Natural Language Processing Techniques
- speechocean762: An Open-Source Non-native English Speech Corpus For Pronunciation Assessment
2021/04/03 by Junbo Zhang, Zhang, Junbo, Zhiwen Zhang +15 · 13 citations
Computer Science · Psychology · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Phonetics and Phonology Research #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and dialogue systems #electronic engineering #information engineering
- Towards Universal Speech Discrete Tokens: A Case Study for ASR and TTS
2023/09/14 by Yifan Yang, Feiyu Shen, Yang, Yifan +11 · 9 citations
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and dialogue systems #Topic Modeling #electronic engineering #information engineering
- PromptASR for contextualized ASR with controllable style
2023/09/14 by Yang, Xiaoyu, Kang, Wei, Yao, Zengwei +5 · 8 citations
#Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
- Pruned RNN-T for fast, memory-efficient ASR training
2022/06/23 by Fangjun Kuang, Kuang, Fangjun, Liyong Guo +10 · 6 citations
Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Natural Language Processing Techniques #Speech Recognition and Synthesis #Speech and dialogue systems #electronic engineering #information engineering
- Lhotse: a speech data representation library for the modern deep learning ecosystem
2021/10/25 by Żelasko, Piotr, Povey, Daniel, Trmal, Jan "Yenda" +1 · 5 citations
#Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
- DOVER-Lap: A Method for Combining Overlap-aware Diarization Outputs
2020/11/03 by Raj, Desh, Garcia-Perera, Leibny Paola, Huang, Zili +4 · 4 citations
#Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
- Improving Neural Biasing for Contextual Speech Recognition by Early Context Injection and Text Perturbation
2024/07/14 by Ruizhe Huang, Mahsa Yarmohammadi, Huang, Ruizhe +5 · 8 citations
Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching
2025/06/16 by Zhu, Han, Kang, Wei, Yao, Zengwei +6 · 11 citations
#Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
- Less Peaky and More Accurate CTC Forced Alignment by Label Priors
2024/04/22 by Ruizhe Huang, Huang, Ruizhe, Xiaohui Zhang +21 · 4 citations
Computer Science · Engineering · #Advanced Numerical Analysis Techniques #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Handwritten Text Recognition Techniques #Machine Learning (cs.LG) #electronic engineering #information engineering
- Krylov Subspace Descent for Deep Learning
2011/11/18 by Vinyals, Oriol, Povey, Daniel · 1 citation
#FOS: Computer and information sciences #FOS: Mathematics #Machine Learning (stat.ML) #Optimization and Control (math.OC)
- k2SSL: A Faster and Better Framework for Self-Supervised Speech Representation Learning
2024/11/26 by Yifan Yang, Jianheng Zhuo, Yang, Yifan +20 · 4 citations
Computer Science · #Speech Recognition and Synthesis #Music and Audio Processing #Speech and Audio Processing
- Predicting Multi-Codebook Vector Quantization Indexes for Knowledge Distillation
2022/10/31 by Liyong Guo, Xiaoyu Yang, Guo, Liyong +20 · 2 citations
Computer Science · #Speech Recognition and Synthesis #Speech and Audio Processing #Music and Audio Processing
- CR-CTC: Consistency regularization on CTC for improved speech recognition
2024/10/07 by Yao, Zengwei, Kang, Wei, Yang, Xiaoyu +7 · 4 citations
#Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Sound (cs.SD) #electronic engineering #information engineering
- SURT 2.0: Advances in Transducer-based Multi-talker Speech Recognition
2023/06/18 by Raj, Desh, Povey, Daniel, Khudanpur, Sanjeev · 2 citations
#Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
- GPU-Accelerated Viterbi Exact Lattice Decoder for Batched Online and Offline Speech Recognition
2019/10/22 by Braun, Hugo, Luitjens, Justin, Leary, Ryan +2 · 1 citation
#Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #electronic engineering #information engineering
- Speaker Diarization with Region Proposal Network
2020/02/14 by Zili Huang, Shinji Watanabe, Huang, Zili +11 · 1 citation
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- Learning from Flawed Data: Weakly Supervised Automatic Speech Recognition
2023/09/26 by Dongji Gao, Gao, Dongji, Xu H +9 · 2 citations
Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Natural Language Processing Techniques #Speech Recognition and Synthesis #Topic Modeling #electronic engineering #information engineering
- Bypass Temporal Classification: Weakly Supervised Automatic Speech Recognition with Imperfect Transcripts
2023/06/01 by Gao, Dongji, Wiesner, Matthew, Xu, Hainan +3 · 2 citations
#Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Sound (cs.SD) #electronic engineering #information engineering
- LibriheavyMix: A 20,000-Hour Dataset for Single-Channel Reverberant Multi-Talker Speech Separation, ASR and Speaker Diarization
2024/09/01 by Zengrui Jin, Jin, Zengrui, Yifan Yang +22 · 2 citations
Computer Science · Health Professions · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Infant Health and Development #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- Fast and parallel decoding for transducer
2022/10/31 by Wei Kang, Liyong Guo, Kang, Wei +14 · 1 citation
Biochemistry, Genetics and Molecular Biology · Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #DNA and Biological Computing #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Network Packet Processing and Optimization #Sound (cs.SD) #Speech Recognition and Synthesis #electronic engineering #information engineering
- Blank-regularized CTC for Frame Skipping in Neural Transducer
2023/05/19 by Yifan Yang, Xiaoyu Yang, Yang, Yifan +14 · 1 citation
Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Neural Networks and Applications #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- Delay-penalized CTC implemented based on Finite State Transducer
2023/05/19 by Yao, Zengwei, Kang, Wei, Kuang, Fangjun +5 · 1 citation
#Audio and Speech Processing (eess.AS) #FOS: Electrical engineering #electronic engineering #information engineering
- ZipVoice-Dialog: Non-Autoregressive Spoken Dialogue Generation with Flow Matching
2025/07/12 by Zhu, Han, Kang, Wei, Guo, Liyong +10 · 4 citations
#Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #electronic engineering #information engineering
- Alternative Pseudo-Labeling for Semi-Supervised Automatic Speech Recognition
2023/08/12 by Zhu Han, Zhu, Han, Dongji Gao +9 · 1 citation
Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- GPU-accelerated Guided Source Separation for Meeting Transcription
2022/12/10 by Raj, Desh, Povey, Daniel, Khudanpur, Sanjeev · 1 citation
#Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering