vix.ing · top · new · best · stats · spec

Shi, Mohan

  1. Semantic VAD: Low-Latency Voice Activity Detection for Speech Interaction
    2023/05/21 by Shi, Mohan, Shu, Yuchun, Zuo, Lingyun +4 · 3 citations
    #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
  2. The second multi-channel multi-party meeting transcription challenge (M2MeT) 2.0): A benchmark for speaker-attributed ASR
    2023/09/24 by Yuhao Liang, Mohan Shi, Liang, Yuhao +24 · 3 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Topic Modeling #electronic engineering #information engineering
  3. Advancing Multi-talker ASR Performance with Large Language Models
    2024/08/30 by Mohan Shi, Zengrui Jin, Shi, Mohan +14 · 4 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Speech Recognition and Synthesis #Speech and dialogue systems #electronic engineering #information engineering
  4. A Comparative Study on Multichannel Speaker-Attributed Automatic Speech Recognition in Multi-party Meetings
    2022/11/01 by Shi, Mohan, Zhang, Jie, Du, Zhihao +4 · 2 citations
    #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
  5. CASA-ASR: Context-Aware Speaker-Attributed ASR
    2023/05/21 by Shi, Mohan, Du, Zhihao, Chen, Qian +5 · 2 citations
    #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
  6. LibriheavyMix: A 20,000-Hour Dataset for Single-Channel Reverberant Multi-Talker Speech Separation, ASR and Speaker Diarization
    2024/09/01 by Zengrui Jin, Jin, Zengrui, Yifan Yang +22 · 2 citations
    Computer Science · Health Professions · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Infant Health and Development #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering