vix.ing · top · new · best · stats · spec

Povey, Daniel

  1. MUSAN: A Music, Speech, and Noise Corpus
    2015/10/28 by Snyder, David, Chen, Guoguo, Povey, Daniel · 77 citations
    #FOS: Computer and information sciences #Sound (cs.SD)
  2. CHiME-6 Challenge:Tackling Multispeaker Speech Recognition for Unsegmented Recordings
    2020/04/20 by Shinji Watanabe, Watanabe, Shinji, Michael Mandel +39 · 27 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  3. Libriheavy: a 50,000 hours ASR corpus with punctuation casing and context
    2023/09/15 by Wei Kang, Kang, Wei, Xiaoyu Yang +12 · 35 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  4. Zipformer: A faster and better encoder for automatic speech recognition
    2023/10/17 by Zengwei Yao, Liyong Guo, Yao, Zengwei +14 · 24 citations
    Computer Science · #Speech Recognition and Synthesis #Music and Audio Processing #Natural Language Processing Techniques
  5. speechocean762: An Open-Source Non-native English Speech Corpus For Pronunciation Assessment
    2021/04/03 by Junbo Zhang, Zhang, Junbo, Zhiwen Zhang +15 · 13 citations
    Computer Science · Psychology · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Phonetics and Phonology Research #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and dialogue systems #electronic engineering #information engineering
  6. Towards Universal Speech Discrete Tokens: A Case Study for ASR and TTS
    2023/09/14 by Yifan Yang, Feiyu Shen, Yang, Yifan +11 · 9 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and dialogue systems #Topic Modeling #electronic engineering #information engineering
  7. PromptASR for contextualized ASR with controllable style
    2023/09/14 by Yang, Xiaoyu, Kang, Wei, Yao, Zengwei +5 · 8 citations
    #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
  8. Pruned RNN-T for fast, memory-efficient ASR training
    2022/06/23 by Fangjun Kuang, Kuang, Fangjun, Liyong Guo +10 · 6 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Natural Language Processing Techniques #Speech Recognition and Synthesis #Speech and dialogue systems #electronic engineering #information engineering
  9. Lhotse: a speech data representation library for the modern deep learning ecosystem
    2021/10/25 by Żelasko, Piotr, Povey, Daniel, Trmal, Jan "Yenda" +1 · 5 citations
    #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
  10. DOVER-Lap: A Method for Combining Overlap-aware Diarization Outputs
    2020/11/03 by Raj, Desh, Garcia-Perera, Leibny Paola, Huang, Zili +4 · 4 citations
    #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
  11. Improving Neural Biasing for Contextual Speech Recognition by Early Context Injection and Text Perturbation
    2024/07/14 by Ruizhe Huang, Mahsa Yarmohammadi, Huang, Ruizhe +5 · 8 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  12. ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching
    2025/06/16 by Zhu, Han, Kang, Wei, Yao, Zengwei +6 · 11 citations
    #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
  13. Less Peaky and More Accurate CTC Forced Alignment by Label Priors
    2024/04/22 by Ruizhe Huang, Huang, Ruizhe, Xiaohui Zhang +21 · 4 citations
    Computer Science · Engineering · #Advanced Numerical Analysis Techniques #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Handwritten Text Recognition Techniques #Machine Learning (cs.LG) #electronic engineering #information engineering
  14. Krylov Subspace Descent for Deep Learning
    2011/11/18 by Vinyals, Oriol, Povey, Daniel · 1 citation
    #FOS: Computer and information sciences #FOS: Mathematics #Machine Learning (stat.ML) #Optimization and Control (math.OC)
  15. k2SSL: A Faster and Better Framework for Self-Supervised Speech Representation Learning
    2024/11/26 by Yifan Yang, Jianheng Zhuo, Yang, Yifan +20 · 4 citations
    Computer Science · #Speech Recognition and Synthesis #Music and Audio Processing #Speech and Audio Processing
  16. Predicting Multi-Codebook Vector Quantization Indexes for Knowledge Distillation
    2022/10/31 by Liyong Guo, Xiaoyu Yang, Guo, Liyong +20 · 2 citations
    Computer Science · #Speech Recognition and Synthesis #Speech and Audio Processing #Music and Audio Processing
  17. CR-CTC: Consistency regularization on CTC for improved speech recognition
    2024/10/07 by Yao, Zengwei, Kang, Wei, Yang, Xiaoyu +7 · 4 citations
    #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Sound (cs.SD) #electronic engineering #information engineering
  18. SURT 2.0: Advances in Transducer-based Multi-talker Speech Recognition
    2023/06/18 by Raj, Desh, Povey, Daniel, Khudanpur, Sanjeev · 2 citations
    #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
  19. GPU-Accelerated Viterbi Exact Lattice Decoder for Batched Online and Offline Speech Recognition
    2019/10/22 by Braun, Hugo, Luitjens, Justin, Leary, Ryan +2 · 1 citation
    #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #electronic engineering #information engineering
  20. Speaker Diarization with Region Proposal Network
    2020/02/14 by Zili Huang, Shinji Watanabe, Huang, Zili +11 · 1 citation
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  21. Learning from Flawed Data: Weakly Supervised Automatic Speech Recognition
    2023/09/26 by Dongji Gao, Gao, Dongji, Xu H +9 · 2 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Natural Language Processing Techniques #Speech Recognition and Synthesis #Topic Modeling #electronic engineering #information engineering
  22. Bypass Temporal Classification: Weakly Supervised Automatic Speech Recognition with Imperfect Transcripts
    2023/06/01 by Gao, Dongji, Wiesner, Matthew, Xu, Hainan +3 · 2 citations
    #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Sound (cs.SD) #electronic engineering #information engineering
  23. LibriheavyMix: A 20,000-Hour Dataset for Single-Channel Reverberant Multi-Talker Speech Separation, ASR and Speaker Diarization
    2024/09/01 by Zengrui Jin, Jin, Zengrui, Yifan Yang +22 · 2 citations
    Computer Science · Health Professions · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Infant Health and Development #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  24. Fast and parallel decoding for transducer
    2022/10/31 by Wei Kang, Liyong Guo, Kang, Wei +14 · 1 citation
    Biochemistry, Genetics and Molecular Biology · Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #DNA and Biological Computing #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Network Packet Processing and Optimization #Sound (cs.SD) #Speech Recognition and Synthesis #electronic engineering #information engineering
  25. Blank-regularized CTC for Frame Skipping in Neural Transducer
    2023/05/19 by Yifan Yang, Xiaoyu Yang, Yang, Yifan +14 · 1 citation
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Neural Networks and Applications #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  26. Delay-penalized CTC implemented based on Finite State Transducer
    2023/05/19 by Yao, Zengwei, Kang, Wei, Kuang, Fangjun +5 · 1 citation
    #Audio and Speech Processing (eess.AS) #FOS: Electrical engineering #electronic engineering #information engineering
  27. ZipVoice-Dialog: Non-Autoregressive Spoken Dialogue Generation with Flow Matching
    2025/07/12 by Zhu, Han, Kang, Wei, Guo, Liyong +10 · 4 citations
    #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #electronic engineering #information engineering
  28. Alternative Pseudo-Labeling for Semi-Supervised Automatic Speech Recognition
    2023/08/12 by Zhu Han, Zhu, Han, Dongji Gao +9 · 1 citation
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  29. GPU-accelerated Guided Source Separation for Meeting Transcription
    2022/12/10 by Raj, Desh, Povey, Daniel, Khudanpur, Sanjeev · 1 citation
    #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering