vix.ing · top · new · best · stats · spec

Chen, Yafeng

  1. CAM++: A Fast and Efficient Network for Speaker Verification Using Context-Aware Masking
    2023/03/01 by Hui Wang, Siqi Zheng, Wang, Hui +7 · 35 citations
    Computer Science · #Speech Recognition and Synthesis #Speech and Audio Processing #Music and Audio Processing
  2. CosyVoice 3: Towards In-the-wild Speech Generation via Scaling-up and Post-training
    2025/05/23 by Zhihao Du, Du, Zhihao, Changfeng Gao +41 · 44 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #Speech Recognition and Synthesis #electronic engineering #information engineering
  3. MinMo: A Multimodal Large Language Model for Seamless Voice Interaction
    2025/01/10 by Chen, Qian, Chen, Yafeng, Chen, Yanni +33 · 25 citations
    #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Human-Computer Interaction (cs.HC) #Sound (cs.SD) #electronic engineering #information engineering
  4. An Enhanced Res2Net with Local and Global Feature Fusion for Speaker Verification
    2023/05/22 by Yafeng Chen, Siqi Zheng, Chen, Yafeng +9 · 5 citations
    Computer Science · #Speech Recognition and Synthesis #Speech and Audio Processing #Music and Audio Processing
  5. 3D-Speaker-Toolkit: An Open-Source Toolkit for Multimodal Speaker Verification and Diarization
    2024/03/29 by Yafeng Chen, Siqi Zheng, Chen, Yafeng +17 · 7 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Electrical engineering #Signal Processing (eess.SP) #Speech Recognition and Synthesis #electronic engineering #information engineering
  6. Exploring Text-Queried Sound Event Detection with Audio Source Separation
    2024/09/20 by Han Yin, Yin, Han, Jisheng Bai +15 · 8 citations
    Computer Science · #Advanced Text Analysis Techniques #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #electronic engineering #information engineering
  7. ERes2NetV2: Boosting Short-Duration Speaker Verification Performance with Computational Efficiency
    2024/06/04 by Yafeng Chen, Siqi Zheng, Chen, Yafeng +11 · 6 citations
    Computer Science · #Speech Recognition and Synthesis #Speech and Audio Processing #Natural Language Processing Techniques
  8. 3D-Speaker: A Large-Scale Multi-Device, Multi-Distance, and Multi-Dialect Corpus for Speech Representation Disentanglement
    2023/06/27 by Siqi Zheng, Luyao Cheng, Zheng, Siqi +7 · 3 citations
    Computer Science · #Speech Recognition and Synthesis #Speech and Audio Processing #Music and Audio Processing
  9. Graph Convolutional Network Based Semi-Supervised Learning on Multi-Speaker Meeting Data
    2022/04/25 by Tong, Fuchuan, Zheng, Siqi, Zhang, Min +4 · 1 citation
    #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
  10. Integrating Audio, Visual, and Semantic Information for Enhanced Multimodal Speaker Diarization
    2024/08/22 by Cheng, Luyao, Wang, Hui, Zheng, Siqi +5 · 2 citations
    #Audio and Speech Processing (eess.AS) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Sound (cs.SD) #electronic engineering #information engineering
  11. Pushing the limits of self-supervised speaker verification using regularized distillation framework
    2022/11/08 by Chen, Yafeng, Zheng, Siqi, Wang, Hui +2 · 1 citation
    #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
  12. SpeakerLM: End-to-End Versatile Speaker Diarization and Recognition with Multimodal Large Language Models
    2025/08/08 by Yin, Han, Chen, Yafeng, Deng, Chong +6 · 4 citations
    #Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Sound (cs.SD)