Gao, Changfeng
- CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models
2024/12/13 by Zhihao Du, Du, Zhihao, Yuxuan Wang +35 · 180 citations
Computer Science · #Speech Recognition and Synthesis #Speech and Audio Processing #Speech and dialogue systems
- FunAudioLLM: Voice Understanding and Generation Foundation Models for Natural Interaction Between Humans and LLMs
2024/07/04 by An, Keyu, Chen, Qian, Deng, Chong +30 · 54 citations
Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #Speech and dialogue systems #electronic engineering #information engineering
- CosyVoice 3: Towards In-the-wild Speech Generation via Scaling-up and Post-training
2025/05/23 by Zhihao Du, Du, Zhihao, Changfeng Gao +41 · 58 citations
Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #Speech Recognition and Synthesis #electronic engineering #information engineering
- MinMo: A Multimodal Large Language Model for Seamless Voice Interaction
2025/01/10 by Chen, Qian, Chen, Yafeng, Chen, Yanni +33 · 47 citations
#Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Human-Computer Interaction (cs.HC) #Sound (cs.SD) #electronic engineering #information engineering
- Differentiable Reward Optimization for LLM based TTS system
2025/07/08 by Changfeng Gao, Zhihao Du, Gao, Changfeng +3 · 8 citations
Computer Science · Psychology · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Phonetics and Phonology Research #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- Transformer-based Online CTC/attention End-to-End Speech Recognition Architecture
2020/01/15 by Haoran Miao, Gaofeng Cheng, Miao, Haoran +7 · 1 citation
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Music and Audio Processing #Natural Language Processing Techniques #Neural and Evolutionary Computing (cs.NE) #Sound (cs.SD) #Speech Recognition and Synthesis #electronic engineering #information engineering
- Fun-ASR Technical Report
2025/09/15 by Keyu An, Yanni Chen, An, Keyu +62 · 3 citations
Computer Science · #Speech Recognition and Synthesis #Speech and Audio Processing #Music and Audio Processing
- Explore the Reinforcement Learning for the LLM based ASR and TTS system
2025/09/23 by Changfeng Gao, Yabin Li, Gao, Changfeng +11 · 1 citation
Computer Science · Psychology · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Emotion and Mood Recognition #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and dialogue systems #electronic engineering #information engineering