He, Jinzheng
- Qwen2 Technical Report
2024/07/15 by Yang, An, Yang, Baosong, Hui, Binyuan +59 · 346 citations
#Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences
- Qwen2-Audio Technical Report
2024/07/15 by Yunfei Chu, Chu, Yunfei, Jin Xu +21 · 1 voice · 155 citations
Computer Science · Engineering · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Music Technology and Sound Studies #cs.CL #cs.LG #eess.AS #electronic engineering #information engineering
- Qwen2.5-Omni Technical Report
2025/03/26 by Jin Xu, Zihan Guo, Xu, Jin +22 · 235 citations
Engineering · #Embedded Systems and FPGA Design
- Real3D-Portrait: One-shot Realistic 3D Talking Portrait Synthesis
2024/01/16 by Ye, Zhenhui, Zhong, Tianyun, Ren, Yi +11 · 19 citations
#Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences
- PopMAG: Pop Music Accompaniment Generation
2020/08/18 by Ren, Yi, He, Jinzheng, Tan, Xu +3 · 7 citations
#Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Information Retrieval (cs.IR) #Machine Learning (cs.LG) #Sound (cs.SD) #electronic engineering #information engineering
- GTSinger: A Global Multi-Technique Singing Corpus with Realistic Music Scores for All Singing Tasks
2024/09/20 by Zhang, Yu, Pan, Changhao, Guo, Wenxiang +15 · 15 citations
#Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
- Qwen3-Omni Technical Report
2025/09/22 by Xu Jin, Zhifang Guo, Xu, Jin +68 · 65 citations
Computer Science · #Speech and Audio Processing #Generative Adversarial Networks and Image Synthesis #Multimodal Machine Learning Applications
- GeneFace++: Generalized and Stable Real-Time Audio-Driven 3D Talking Face Generation
2023/05/01 by Zhenhui Ye, Jinzheng He, Ye, Zhenhui +17 · 8 citations
Computer Science · #Face recognition and analysis #Generative Adversarial Networks and Image Synthesis #Speech and Audio Processing
- RMSSinger: Realistic-Music-Score based Singing Voice Synthesis
2023/05/18 by Jinzheng He, He, Jinzheng, Jinglin Liu +11 · 7 citations
Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- Mega-TTS 2: Boosting Prompting Mechanisms for Zero-Shot Speech Synthesis
2023/07/14 by Ziyue Karen Jiang, Jiang, Ziyue, Jinglin Liu +21 · 7 citations
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- TCSinger: Zero-Shot Singing Voice Synthesis with Style Transfer and Multi-Level Style Control
2024/09/24 by Zhang, Yu, Jiang, Ziyue, Li, Ruiqi +5 · 8 citations
#Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
- MimicTalk: Mimicking a personalized and expressive 3D talking face in minutes
2024/10/09 by Ye, Zhenhui, Zhong, Tianyun, Ren, Yi +10 · 6 citations
#Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences
- WavRAG: Audio-Integrated Retrieval Augmented Generation for Spoken Dialogue Models
2025/02/20 by Y.W. Chen, Chen, Yifu, Shengpeng Ji +13 · 8 citations
Computer Science · #Speech and dialogue systems #Topic Modeling #Natural Language Processing Techniques
- CLAPSpeech: Learning Prosody from Text Context with Contrastive Language-Audio Pre-training
2023/05/18 by Ye, Zhenhui, Huang, Rongjie, Ren, Yi +5 · 3 citations
#Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
- ViT-TTS: Visual Text-to-Speech with Scalable Diffusion Transformer
2023/05/22 by Liu, Huadai, Huang, Rongjie, Lin, Xuan +5 · 2 citations
#Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
- WavReward: Spoken Dialogue Models With Generalist Reward Evaluators
2025/05/14 by Shengpeng Ji, Ji, Shengpeng, Yangzhuo Li +22 · 3 citations
Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Multi-Agent Systems and Negotiation #Multimedia (cs.MM) #Sound (cs.SD) #Speech and dialogue systems #Topic Modeling #electronic engineering #information engineering
- Omni-Captioner: Data Pipeline, Models, and Benchmark for Omni Detailed Perception
2025/10/14 by Ma, Ziyang, Xu, Ruiyang, Xing, Zhenghao +9 · 2 citations
#Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Multimedia (cs.MM) #Sound (cs.SD)