vix.ing · top · new · best · stats · spec

Raj, Desh

  1. CHiME-6 Challenge:Tackling Multispeaker Speech Recognition for Unsegmented Recordings
    2020/04/20 by Shinji Watanabe, Watanabe, Shinji, Michael Mandel +39 · 22 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  2. Integration of speech separation, diarization, and recognition for multi-speaker meetings: System description, comparison, and analysis
    2020/11/03 by Raj, Desh, Denisov, Pavel, Chen, Zhuo +11 · 11 citations
    #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
  3. DOVER-Lap: A Method for Combining Overlap-aware Diarization Outputs
    2020/11/03 by Raj, Desh, Garcia-Perera, Leibny Paola, Huang, Zili +4 · 4 citations
    #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
  4. The CHiME-7 DASR Challenge: Distant Meeting Transcription with Multiple Devices in Diverse Scenarios
    2023/06/23 by Cornell, Samuele, Wiesner, Matthew, Watanabe, Shinji +8 · 5 citations
    #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
  5. Adapting self-supervised models to multi-talker speech recognition using speaker embeddings
    2022/11/01 by Zili Huang, Huang, Zili, Desh Raj +5 · 3 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and dialogue systems #electronic engineering #information engineering
  6. The Hitachi-JHU DIHARD III System: Competitive End-to-End Neural Diarization and X-Vector Clustering Systems Combined by DOVER-Lap
    2021/02/02 by Horiguchi, Shota, Yalta, Nelson, Garcia, Paola +7 · 2 citations
    #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
  7. SURT 2.0: Advances in Transducer-based Multi-talker Speech Recognition
    2023/06/18 by Raj, Desh, Povey, Daniel, Khudanpur, Sanjeev · 2 citations
    #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
  8. Learning from Flawed Data: Weakly Supervised Automatic Speech Recognition
    2023/09/26 by Dongji Gao, Xu H, Gao, Dongji +9 · 2 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Natural Language Processing Techniques #Speech Recognition and Synthesis #Topic Modeling #electronic engineering #information engineering
  9. Target-speaker Voice Activity Detection with Improved I-Vector Estimation for Unknown Number of Speaker
    2021/08/07 by He, Maokui, Raj, Desh, Huang, Zili +3 · 1 citation
    #Audio and Speech Processing (eess.AS) #FOS: Electrical engineering #electronic engineering #information engineering
  10. Continuous Streaming Multi-Talker ASR with Dual-path Transducers
    2021/09/17 by Raj, Desh, Lu, Liang, Chen, Zhuo +2 · 1 citation
    #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
  11. Low-Latency Speech Separation Guided Diarization for Telephone Conversations
    2022/04/05 by Giovanni Morrone, Samuele Cornell, Morrone, Giovanni +10 · 1 citation
    Computer Science · Psychology · #Audio and Speech Processing (eess.AS) #FOS: Electrical engineering #Phonetics and Phonology Research #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  12. Training dynamic models using early exits for automatic speech recognition on resource-constrained devices
    2023/09/18 by Wright, George August, Cappellazzo, Umberto, Zaiem, Salah +5 · 1 citation
    #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
  13. Updated Corpora and Benchmarks for Long-Form Speech Recognition
    2023/09/26 by Fox, Jennifer Drexler, Raj, Desh, Delworth, Natalie +3 · 1 citation
    #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
  14. GPU-accelerated Guided Source Separation for Meeting Transcription
    2022/12/10 by Raj, Desh, Povey, Daniel, Khudanpur, Sanjeev · 1 citation
    #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
  15. M-BEST-RQ: A Multi-Channel Speech Foundation Model for Smart Glasses
    2024/09/17 by Yang, Yufeng, Raj, Desh, Lin, Ju +8 · 1 citation
    #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering