Du, Chenpeng
- AniTalker: Animate Vivid and Diverse Talking Faces through Identity-Decoupled Facial Motion Encoding
2024/05/06 by Tao Liu, Liu, Tao, Feilong Chen +11 · 21 citations
Computer Science · #Artificial Intelligence (cs.AI) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Face recognition and analysis
- VoiceFlow: Efficient Text-to-Speech with Rectified Flow Matching
2023/09/10 by Guo, Yiwei, Du, Chenpeng, Ma, Ziyang +2 · 13 citations
#Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Human-Computer Interaction (cs.HC) #Sound (cs.SD) #electronic engineering #information engineering
- Language Model Can Listen While Speaking
2024/08/05 by Ziyang Ma, Ma, Ziyang, Song, Yakun +12 · 17 citations
Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Human-Computer Interaction (cs.HC) #Natural Language Processing Techniques #Sound (cs.SD) #Speech and dialogue systems #electronic engineering #information engineering
- DiTAR: Diffusion Transformer Autoregressive Modeling for Speech Generation
2025/02/06 by Jia, Dongya, Chen, Zhuo, Chen, Jiawei +8 · 23 citations
#Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Sound (cs.SD) #electronic engineering #information engineering
- Recent Advances in Discrete Speech Tokens: A Review
2025/02/10 by Yiwei Guo, Guo, Yiwei, Zhihan Li +16 · 24 citations
Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Internet Traffic Analysis and Secure E-voting #Multimedia (cs.MM) #Signal Processing (eess.SP) #Sound (cs.SD) #Speech Recognition and Synthesis #electronic engineering #information engineering #semigroups and automata theory
- EmoDiff: Intensity Controllable Emotional Text-to-Speech with Soft-Label Guidance
2022/11/17 by Yiwei Guo, Chenpeng Du, Guo, Yiwei +5 · 8 citations
Computer Science · Psychology · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Human-Computer Interaction (cs.HC) #Machine Learning (cs.LG) #Phonetics and Phonology Research #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- Towards Universal Speech Discrete Tokens: A Case Study for ASR and TTS
2023/09/14 by Yifan Yang, Yang, Yifan, Feiyu Shen +11 · 9 citations
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and dialogue systems #Topic Modeling #electronic engineering #information engineering
- LSCodec: Low-Bitrate and Speaker-Decoupled Discrete Speech Codec
2024/10/21 by Yiwei Guo, Zhihan Li, Guo, Yiwei +9 · 9 citations
Computer Science · #Advanced Data Compression Techniques #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- DiffDub: Person-generic Visual Dubbing Using Inpainting Renderer with Diffusion Auto-encoder
2023/11/03 by Tao Liu, Liu, Tao, Chenpeng Du +7 · 3 citations
Computer Science · #Artificial Intelligence (cs.AI) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Face recognition and analysis #Generative Adversarial Networks and Image Synthesis #Multimodal Machine Learning Applications
- Attention-Constrained Inference for Robust Decoder-Only Text-to-Speech
2024/04/30 by Hankun Wang, Chenpeng Du, Wang, Hankun +9 · 2 citations
Computer Science · #Speech Recognition and Synthesis
- GSTalker: Real-time Audio-Driven Talking Face Generation via Deformable Gaussian Splatting
2024/04/29 by Chen, Bo, Hu, Shoukang, Chen, Qi +4 · 2 citations
#Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences
- Unsupervised word-level prosody tagging for controllable speech synthesis
2022/02/15 by Guo, Yiwei, Du, Chenpeng, Yu, Kai · 1 citation
#Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Sound (cs.SD) #electronic engineering #information engineering
- DSE-TTS: Dual Speaker Embedding for Cross-Lingual Text-to-Speech
2023/06/25 by Sen Liu, Yiwei Guo, Liu, Sen +7 · 1 citation
Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Topic Modeling #electronic engineering #information engineering
- Acoustic BPE for Speech Generation with Discrete Tokens
2023/10/23 by Feiyu Shen, Shen, Feiyu, Yiwei Guo +7 · 1 citation
Computer Science · #Speech Recognition and Synthesis #Speech and dialogue systems #Topic Modeling
- MagiCodec: Simple Masked Gaussian-Injected Codec for High-Fidelity Reconstruction and Generation
2025/05/31 by Song, Yakun, Chen, Jiawei, Zhuang, Xiaobin +9 · 3 citations
#Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Sound (cs.SD) #electronic engineering #information engineering
- vec2wav 2.0: Advancing Voice Conversion via Discrete Token Vocoders
2024/09/03 by Yiwei Guo, Guo, Yiwei, Zhihan Li +13 · 1 citation
Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- Towards Reliable Large Audio Language Model
2025/05/25 by Ziyang Ma, Ma, Ziyang, Xiquan Li +17 · 2 citations
Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Human-Computer Interaction (cs.HC) #Multimedia (cs.MM) #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- Why Do Speech Language Models Fail to Generate Semantically Coherent Outputs? A Modality Evolving Perspective
2024/12/22 by Wang, Hankun, Wang, Haoran, Guo, Yiwei +4 · 1 citation
#Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering