vix.ing · top · new · best · stats · spec

He, Jinzheng

  1. Qwen2 Technical Report
    2024/07/15 by Yang, An, Yang, Baosong, Hui, Binyuan +59 · 346 citations
    #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences
  2. Qwen2-Audio Technical Report
    2024/07/15 by Yunfei Chu, Chu, Yunfei, Jin Xu +21 · 1 voice · 155 citations
    Computer Science · Engineering · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Music Technology and Sound Studies #cs.CL #cs.LG #eess.AS #electronic engineering #information engineering
  3. Qwen2.5-Omni Technical Report
    2025/03/26 by Jin Xu, Zihan Guo, Xu, Jin +22 · 235 citations
    Engineering · #Embedded Systems and FPGA Design
  4. Real3D-Portrait: One-shot Realistic 3D Talking Portrait Synthesis
    2024/01/16 by Ye, Zhenhui, Zhong, Tianyun, Ren, Yi +11 · 19 citations
    #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences
  5. PopMAG: Pop Music Accompaniment Generation
    2020/08/18 by Ren, Yi, He, Jinzheng, Tan, Xu +3 · 7 citations
    #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Information Retrieval (cs.IR) #Machine Learning (cs.LG) #Sound (cs.SD) #electronic engineering #information engineering
  6. GTSinger: A Global Multi-Technique Singing Corpus with Realistic Music Scores for All Singing Tasks
    2024/09/20 by Zhang, Yu, Pan, Changhao, Guo, Wenxiang +15 · 15 citations
    #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
  7. Qwen3-Omni Technical Report
    2025/09/22 by Xu Jin, Zhifang Guo, Xu, Jin +68 · 65 citations
    Computer Science · #Speech and Audio Processing #Generative Adversarial Networks and Image Synthesis #Multimodal Machine Learning Applications
  8. GeneFace++: Generalized and Stable Real-Time Audio-Driven 3D Talking Face Generation
    2023/05/01 by Zhenhui Ye, Jinzheng He, Ye, Zhenhui +17 · 8 citations
    Computer Science · #Face recognition and analysis #Generative Adversarial Networks and Image Synthesis #Speech and Audio Processing
  9. RMSSinger: Realistic-Music-Score based Singing Voice Synthesis
    2023/05/18 by Jinzheng He, He, Jinzheng, Jinglin Liu +11 · 7 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  10. Mega-TTS 2: Boosting Prompting Mechanisms for Zero-Shot Speech Synthesis
    2023/07/14 by Ziyue Karen Jiang, Jiang, Ziyue, Jinglin Liu +21 · 7 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  11. TCSinger: Zero-Shot Singing Voice Synthesis with Style Transfer and Multi-Level Style Control
    2024/09/24 by Zhang, Yu, Jiang, Ziyue, Li, Ruiqi +5 · 8 citations
    #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
  12. MimicTalk: Mimicking a personalized and expressive 3D talking face in minutes
    2024/10/09 by Ye, Zhenhui, Zhong, Tianyun, Ren, Yi +10 · 6 citations
    #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences
  13. WavRAG: Audio-Integrated Retrieval Augmented Generation for Spoken Dialogue Models
    2025/02/20 by Y.W. Chen, Chen, Yifu, Shengpeng Ji +13 · 8 citations
    Computer Science · #Speech and dialogue systems #Topic Modeling #Natural Language Processing Techniques
  14. CLAPSpeech: Learning Prosody from Text Context with Contrastive Language-Audio Pre-training
    2023/05/18 by Ye, Zhenhui, Huang, Rongjie, Ren, Yi +5 · 3 citations
    #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
  15. ViT-TTS: Visual Text-to-Speech with Scalable Diffusion Transformer
    2023/05/22 by Liu, Huadai, Huang, Rongjie, Lin, Xuan +5 · 2 citations
    #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
  16. WavReward: Spoken Dialogue Models With Generalist Reward Evaluators
    2025/05/14 by Shengpeng Ji, Ji, Shengpeng, Yangzhuo Li +22 · 3 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Multi-Agent Systems and Negotiation #Multimedia (cs.MM) #Sound (cs.SD) #Speech and dialogue systems #Topic Modeling #electronic engineering #information engineering
  17. Omni-Captioner: Data Pipeline, Models, and Benchmark for Omni Detailed Perception
    2025/10/14 by Ma, Ziyang, Xu, Ruiyang, Xing, Zhenghao +9 · 2 citations
    #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Multimedia (cs.MM) #Sound (cs.SD)