Chen, Yafeng
- CAM++: A Fast and Efficient Network for Speaker Verification Using Context-Aware Masking
2023/03/01 by Hui Wang, Siqi Zheng, Wang, Hui +7 · 35 citations
Computer Science · #Speech Recognition and Synthesis #Speech and Audio Processing #Music and Audio Processing
- CosyVoice 3: Towards In-the-wild Speech Generation via Scaling-up and Post-training
2025/05/23 by Zhihao Du, Du, Zhihao, Changfeng Gao +41 · 44 citations
Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #Speech Recognition and Synthesis #electronic engineering #information engineering
- MinMo: A Multimodal Large Language Model for Seamless Voice Interaction
2025/01/10 by Chen, Qian, Chen, Yafeng, Chen, Yanni +33 · 25 citations
#Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Human-Computer Interaction (cs.HC) #Sound (cs.SD) #electronic engineering #information engineering
- An Enhanced Res2Net with Local and Global Feature Fusion for Speaker Verification
2023/05/22 by Yafeng Chen, Siqi Zheng, Chen, Yafeng +9 · 5 citations
Computer Science · #Speech Recognition and Synthesis #Speech and Audio Processing #Music and Audio Processing
- 3D-Speaker-Toolkit: An Open-Source Toolkit for Multimodal Speaker Verification and Diarization
2024/03/29 by Yafeng Chen, Siqi Zheng, Chen, Yafeng +17 · 7 citations
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Electrical engineering #Signal Processing (eess.SP) #Speech Recognition and Synthesis #electronic engineering #information engineering
- Exploring Text-Queried Sound Event Detection with Audio Source Separation
2024/09/20 by Han Yin, Yin, Han, Jisheng Bai +15 · 8 citations
Computer Science · #Advanced Text Analysis Techniques #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #electronic engineering #information engineering
- ERes2NetV2: Boosting Short-Duration Speaker Verification Performance with Computational Efficiency
2024/06/04 by Yafeng Chen, Siqi Zheng, Chen, Yafeng +11 · 6 citations
Computer Science · #Speech Recognition and Synthesis #Speech and Audio Processing #Natural Language Processing Techniques
- 3D-Speaker: A Large-Scale Multi-Device, Multi-Distance, and Multi-Dialect Corpus for Speech Representation Disentanglement
2023/06/27 by Siqi Zheng, Luyao Cheng, Zheng, Siqi +7 · 3 citations
Computer Science · #Speech Recognition and Synthesis #Speech and Audio Processing #Music and Audio Processing
- Graph Convolutional Network Based Semi-Supervised Learning on Multi-Speaker Meeting Data
2022/04/25 by Tong, Fuchuan, Zheng, Siqi, Zhang, Min +4 · 1 citation
#Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
- Integrating Audio, Visual, and Semantic Information for Enhanced Multimodal Speaker Diarization
2024/08/22 by Cheng, Luyao, Wang, Hui, Zheng, Siqi +5 · 2 citations
#Audio and Speech Processing (eess.AS) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Sound (cs.SD) #electronic engineering #information engineering
- Pushing the limits of self-supervised speaker verification using regularized distillation framework
2022/11/08 by Chen, Yafeng, Zheng, Siqi, Wang, Hui +2 · 1 citation
#Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
- SpeakerLM: End-to-End Versatile Speaker Diarization and Recognition with Multimodal Large Language Models
2025/08/08 by Yin, Han, Chen, Yafeng, Deng, Chong +6 · 4 citations
#Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Sound (cs.SD)