Zhenhui Ye
- AudioGPT: Understanding and Generating Speech, Music, Sound, and Talking Head
2023/04/25 by Rongjie Huang, Mingze Li, Huang, Rongjie +23 · 1 voice · 46 citations
Computer Science · #Music and Audio Processing #Speech Recognition and Synthesis #Topic Modeling
- GeneFace: Generalized and High-Fidelity Audio-Driven 3D Talking Face Synthesis
2023/01/31 by Zhenhui Ye, Ziyue Karen Jiang, Ye, Zhenhui +9 · 19 citations
Computer Science · #Face recognition and analysis #Generative Adversarial Networks and Image Synthesis #Speech and Audio Processing
- Make-An-Audio 2: Temporal-Enhanced Text-to-Audio Generation
2023/05/29 by Jiawei Huang, Huang, Jiawei, Yi Ren +17 · 17 citations
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Multimedia (cs.MM) #Music Technology and Sound Studies #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #electronic engineering #information engineering
- GeneFace++: Generalized and Stable Real-Time Audio-Driven 3D Talking Face Generation
2023/05/01 by Zhenhui Ye, Jinzheng He, Ye, Zhenhui +17 · 8 citations
Computer Science · #Face recognition and analysis #Generative Adversarial Networks and Image Synthesis #Speech and Audio Processing
- MegaTTS 3: Sparse Alignment Enhanced Latent Diffusion Transformer for Zero-Shot Speech Synthesis
2025/02/26 by Ziyue Karen Jiang, Yi Ren, Jiang, Ziyue +24 · 19 citations
Computer Science · Medicine · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Face recognition and analysis #Machine Learning (cs.LG) #Sound (cs.SD) #Speech Recognition and Synthesis #Voice and Speech Disorders #electronic engineering #information engineering
- RMSSinger: Realistic-Music-Score based Singing Voice Synthesis
2023/05/18 by Jinzheng He, Jinglin Liu, He, Jinzheng +11 · 7 citations
Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- Mega-TTS: Zero-Shot Text-to-Speech at Scale with Intrinsic Inductive Bias
2023/06/06 by Ziyue Karen Jiang, Jiang, Ziyue, Yi Ren +21 · 7 citations
Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- Mega-TTS 2: Boosting Prompting Mechanisms for Zero-Shot Speech Synthesis
2023/07/14 by Ziyue Karen Jiang, Jiang, Ziyue, Jinglin Liu +21 · 7 citations
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- FluentSpeech: Stutter-Oriented Automatic Speech Editing with Context-Aware Diffusion Models
2023/05/23 by Ziyue Karen Jiang, Jiang, Ziyue, Qian Yang +11 · 5 citations
Computer Science · Psychology · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Phonetics and Phonology Research #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- HumanDiT: Pose-Guided Diffusion Transformer for Long-form Human Motion Video Generation
2025/02/07 by Qijun Gan, Yi Ren, Gan, Qijun +15 · 12 citations
Computer Science · Engineering · #Advanced Vision and Imaging #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Human Motion and Animation #Human Pose and Action Recognition
- Make-A-Voice: Unified Voice Synthesis With Discrete Representation
2023/05/30 by Rongjie Huang, Huang, Rongjie, Chunlei Zhang +17 · 5 citations
Computer Science · #Speech Recognition and Synthesis #Music and Audio Processing #Topic Modeling
- FreeBind: Free Lunch in Unified Multimodal Space via Knowledge Fusion
2024/05/08 by Zehan Wang, Ziang Zhang, Wang, Zehan +19 · 5 citations
Computer Science · Chemistry · Materials Science · #Semantic Web and Ontologies #History and advancements in chemistry #Diatoms and Algae Research
- Ada-TTA: Towards Adaptive High-Quality Text-to-Talking Avatar Synthesis
2023/06/06 by Zhenhui Ye, Ye, Zhenhui, Ziyue Karen Jiang +13 · 1 citation
Computer Science · #Face recognition and analysis #Generative Adversarial Networks and Image Synthesis #Speech and Audio Processing