vix.ing · top · new · best · stats · spec

Chunyu Qiang

  1. Kling-Foley: Multimodal Diffusion Transformer for High-Quality Video-to-Audio Generation
    2025/06/24 by Jun Wang, Wang, Jun, Xijuan Zeng +40 · 19 citations
    Computer Science · Social Sciences · #Speech and Audio Processing #Multimedia Communication and Technology
  2. Learning Speech Representation From Contrastive Token-Acoustic Pretraining
    2023/09/01 by Chunyu Qiang, Hao Li, Qiang, Chunyu +11 · 3 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  3. VQ-CTAP: Cross-Modal Fine-Grained Sequence Representation Learning for Speech Processing
    2024/08/11 by Chunyu Qiang, Geng Wang, Qiang, Chunyu +26 · 5 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and dialogue systems #electronic engineering #information engineering
  4. An Initial Investigation of Language Adaptation for TTS Systems under Low-resource Scenarios
    2024/06/13 by Gong Cheng, Gong, Cheng, Erica Cooper +21 · 3 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Multimodal Machine Learning Applications #Speech Recognition and Synthesis #Speech and dialogue systems #electronic engineering #information engineering
  5. Style-Label-Free: Cross-Speaker Style Transfer by Quantized VAE and Speaker-wise Normalization in Speech Synthesis
    2022/12/13 by Chunyu Qiang, Peng Yang, Qiang, Chunyu +7 · 1 citation
    Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  6. ASRRL-TTS: Agile Speaker Representation Reinforcement Learning for Text-to-Speech Speaker Adaptation
    2024/07/07 by Ruibo Fu, Fu, Ruibo, Xin Qi +23 · 1 citation
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering