vix.ing · top · new · best · stats · spec

Gao, Yingming

  1. Auffusion: Leveraging the Power of Diffusion and Large Language Models for Text-to-Audio Generation
    2024/01/02 by Jinlong Xue, Yayue Deng, Xue, Jinlong +5 · 13 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Generative Adversarial Networks and Image Synthesis #Music Technology and Sound Studies #Music and Audio Processing #Sound (cs.SD) #electronic engineering #information engineering
  2. Retrieval Augmented Generation in Prompt-based Text-to-Speech Synthesis with Context-Aware Contrastive Language-Audio Pretraining
    2024/06/06 by Xue, Jinlong, Deng, Yayue, Gao, Yingming +1 · 5 citations
    #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
  3. SPA-SVC: Self-supervised Pitch Augmentation for Singing Voice Conversion
    2024/06/09 by Bingsong Bai, Fengping Wang, Bai, Bingsong +5 · 4 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  4. Enhancing Modal Fusion by Alignment and Label Matching for Multimodal Emotion Recognition
    2024/08/18 by Qifei Li, Li, Qifei, Yingming Gao +7 · 3 citations
    Psychology · Computer Science · #Emotion and Mood Recognition #Human Pose and Action Recognition #Video Surveillance and Tracking Methods
  5. Improving Audio Codec-based Zero-Shot Text-to-Speech Synthesis with Multi-Modal Context and Large Language Model
    2024/06/06 by Jinlong Xue, Xue, Jinlong, Yayue Deng +7 · 2 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #Speech Recognition and Synthesis #electronic engineering #information engineering
  6. Text-Aware End-to-end Mispronunciation Detection and Diagnosis
    2022/06/15 by Linkai Peng, Yingming Gao, Peng, Linkai +9 · 1 citation
    Computer Science · Psychology · #Speech Recognition and Synthesis #Speech and Audio Processing #Phonetics and Phonology Research
  7. CONCSS: Contrastive-based Context Comprehension for Dialogue-appropriate Prosody in Conversational Speech Synthesis
    2023/12/16 by Deng, Yayue, Xue, Jinlong, Jia, Yukang +6 · 1 citation
    #Computation and Language (cs.CL) #FOS: Computer and information sciences #Human-Computer Interaction (cs.HC)
  8. Frame-level emotional state alignment method for speech emotion recognition
    2023/12/27 by Qifei Li, Yingming Gao, Li, Qifei +11 · 1 citation
    Computer Science · Neuroscience · Psychology · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #EEG and Brain-Computer Interfaces #Emotion and Mood Recognition #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #Speech Recognition and Synthesis #electronic engineering #information engineering
  9. HQ-SVC: Towards High-Quality Zero-Shot Singing Voice Conversion in Low-Resource Scenarios
    2025/11/11 by Bingsong Bai, Yanli Geng, Bai, Bingsong +10 · 2 citations
    Computer Science · Medicine · #Speech and Audio Processing #Speech Recognition and Synthesis #Voice and Speech Disorders