Raj, Desh
- CHiME-6 Challenge:Tackling Multispeaker Speech Recognition for Unsegmented Recordings
2020/04/20 by Shinji Watanabe, Watanabe, Shinji, Michael Mandel +39 · 22 citations
Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- Integration of speech separation, diarization, and recognition for multi-speaker meetings: System description, comparison, and analysis
2020/11/03 by Raj, Desh, Denisov, Pavel, Chen, Zhuo +11 · 11 citations
#Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
- DOVER-Lap: A Method for Combining Overlap-aware Diarization Outputs
2020/11/03 by Raj, Desh, Garcia-Perera, Leibny Paola, Huang, Zili +4 · 4 citations
#Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
- The CHiME-7 DASR Challenge: Distant Meeting Transcription with Multiple Devices in Diverse Scenarios
2023/06/23 by Cornell, Samuele, Wiesner, Matthew, Watanabe, Shinji +8 · 5 citations
#Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
- Adapting self-supervised models to multi-talker speech recognition using speaker embeddings
2022/11/01 by Zili Huang, Huang, Zili, Desh Raj +5 · 3 citations
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and dialogue systems #electronic engineering #information engineering
- The Hitachi-JHU DIHARD III System: Competitive End-to-End Neural Diarization and X-Vector Clustering Systems Combined by DOVER-Lap
2021/02/02 by Horiguchi, Shota, Yalta, Nelson, Garcia, Paola +7 · 2 citations
#Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
- SURT 2.0: Advances in Transducer-based Multi-talker Speech Recognition
2023/06/18 by Raj, Desh, Povey, Daniel, Khudanpur, Sanjeev · 2 citations
#Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
- Learning from Flawed Data: Weakly Supervised Automatic Speech Recognition
2023/09/26 by Dongji Gao, Xu H, Gao, Dongji +9 · 2 citations
Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Natural Language Processing Techniques #Speech Recognition and Synthesis #Topic Modeling #electronic engineering #information engineering
- Target-speaker Voice Activity Detection with Improved I-Vector Estimation for Unknown Number of Speaker
2021/08/07 by He, Maokui, Raj, Desh, Huang, Zili +3 · 1 citation
#Audio and Speech Processing (eess.AS) #FOS: Electrical engineering #electronic engineering #information engineering
- Continuous Streaming Multi-Talker ASR with Dual-path Transducers
2021/09/17 by Raj, Desh, Lu, Liang, Chen, Zhuo +2 · 1 citation
#Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
- Low-Latency Speech Separation Guided Diarization for Telephone Conversations
2022/04/05 by Giovanni Morrone, Samuele Cornell, Morrone, Giovanni +10 · 1 citation
Computer Science · Psychology · #Audio and Speech Processing (eess.AS) #FOS: Electrical engineering #Phonetics and Phonology Research #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- Training dynamic models using early exits for automatic speech recognition on resource-constrained devices
2023/09/18 by Wright, George August, Cappellazzo, Umberto, Zaiem, Salah +5 · 1 citation
#Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
- Updated Corpora and Benchmarks for Long-Form Speech Recognition
2023/09/26 by Fox, Jennifer Drexler, Raj, Desh, Delworth, Natalie +3 · 1 citation
#Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
- GPU-accelerated Guided Source Separation for Meeting Transcription
2022/12/10 by Raj, Desh, Povey, Daniel, Khudanpur, Sanjeev · 1 citation
#Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
- M-BEST-RQ: A Multi-Channel Speech Foundation Model for Smart Glasses
2024/09/17 by Yang, Yufeng, Raj, Desh, Lin, Ju +8 · 1 citation
#Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering