Gao, Yingming
- Auffusion: Leveraging the Power of Diffusion and Large Language Models for Text-to-Audio Generation
2024/01/02 by Jinlong Xue, Yayue Deng, Xue, Jinlong +5 · 13 citations
Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Generative Adversarial Networks and Image Synthesis #Music Technology and Sound Studies #Music and Audio Processing #Sound (cs.SD) #electronic engineering #information engineering
- Retrieval Augmented Generation in Prompt-based Text-to-Speech Synthesis with Context-Aware Contrastive Language-Audio Pretraining
2024/06/06 by Xue, Jinlong, Deng, Yayue, Gao, Yingming +1 · 5 citations
#Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
- SPA-SVC: Self-supervised Pitch Augmentation for Singing Voice Conversion
2024/06/09 by Bingsong Bai, Fengping Wang, Bai, Bingsong +5 · 4 citations
Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- Enhancing Modal Fusion by Alignment and Label Matching for Multimodal Emotion Recognition
2024/08/18 by Qifei Li, Li, Qifei, Yingming Gao +7 · 3 citations
Psychology · Computer Science · #Emotion and Mood Recognition #Human Pose and Action Recognition #Video Surveillance and Tracking Methods
- Improving Audio Codec-based Zero-Shot Text-to-Speech Synthesis with Multi-Modal Context and Large Language Model
2024/06/06 by Jinlong Xue, Xue, Jinlong, Yayue Deng +7 · 2 citations
Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #Speech Recognition and Synthesis #electronic engineering #information engineering
- Text-Aware End-to-end Mispronunciation Detection and Diagnosis
2022/06/15 by Linkai Peng, Yingming Gao, Peng, Linkai +9 · 1 citation
Computer Science · Psychology · #Speech Recognition and Synthesis #Speech and Audio Processing #Phonetics and Phonology Research
- CONCSS: Contrastive-based Context Comprehension for Dialogue-appropriate Prosody in Conversational Speech Synthesis
2023/12/16 by Deng, Yayue, Xue, Jinlong, Jia, Yukang +6 · 1 citation
#Computation and Language (cs.CL) #FOS: Computer and information sciences #Human-Computer Interaction (cs.HC)
- Frame-level emotional state alignment method for speech emotion recognition
2023/12/27 by Qifei Li, Yingming Gao, Li, Qifei +11 · 1 citation
Computer Science · Neuroscience · Psychology · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #EEG and Brain-Computer Interfaces #Emotion and Mood Recognition #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #Speech Recognition and Synthesis #electronic engineering #information engineering
- HQ-SVC: Towards High-Quality Zero-Shot Singing Voice Conversion in Low-Resource Scenarios
2025/11/11 by Bingsong Bai, Yanli Geng, Bai, Bingsong +10 · 2 citations
Computer Science · Medicine · #Speech and Audio Processing #Speech Recognition and Synthesis #Voice and Speech Disorders