vix.ing · top · new · best · stats · spec

Jinzheng He

  1. Qwen2-Audio Technical Report
    2024/07/15 by Yunfei Chu, Chu, Yunfei, Jin Xu +21 · 1 voice · 160 citations
    Computer Science · Engineering · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Music Technology and Sound Studies #cs.CL #cs.LG #eess.AS #electronic engineering #information engineering
  2. Qwen3-Omni Technical Report
    2025/09/22 by Jin Xu, Xu Jin, Xu, Jin +78 · 1 voice · 66 citations
    Computer Science · Engineering · #Generative Adversarial Networks and Image Synthesis #Multimodal Machine Learning Applications #Speech and Audio Processing #cs.AI #cs.CL #cs.CV #eess.AS
  3. GeneFace++: Generalized and Stable Real-Time Audio-Driven 3D Talking Face Generation
    2023/05/01 by Zhenhui Ye, Jinzheng He, Ye, Zhenhui +17 · 8 citations
    Computer Science · #Face recognition and analysis #Generative Adversarial Networks and Image Synthesis #Speech and Audio Processing
  4. RMSSinger: Realistic-Music-Score based Singing Voice Synthesis
    2023/05/18 by Jinzheng He, Jinglin Liu, He, Jinzheng +11 · 7 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  5. Mega-TTS 2: Boosting Prompting Mechanisms for Zero-Shot Speech Synthesis
    2023/07/14 by Ziyue Karen Jiang, Jinglin Liu, Jiang, Ziyue +21 · 7 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  6. CLAPSpeech: Learning Prosody from Text Context with Contrastive Language-Audio Pre-training
    2023/05/18 by Zhenhui Ye, Rongjie Huang, Ye, Zhenhui +13 · 4 citations
    Computer Science · Psychology · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Phonetics and Phonology Research #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and dialogue systems #electronic engineering #information engineering
  7. WavRAG: Audio-Integrated Retrieval Augmented Generation for Spoken Dialogue Models
    2025/02/20 by Y.W. Chen, Shengpeng Ji, Chen, Yifu +13 · 8 citations
    Computer Science · #Speech and dialogue systems #Topic Modeling #Natural Language Processing Techniques
  8. WavReward: Spoken Dialogue Models With Generalist Reward Evaluators
    2025/05/14 by Shengpeng Ji, Ji, Shengpeng, Liang, Tianle +22 · 3 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Multi-Agent Systems and Negotiation #Multimedia (cs.MM) #Sound (cs.SD) #Speech and dialogue systems #Topic Modeling #electronic engineering #information engineering