vix.ing · top · new · best · stats · spec

Du, Chenpeng

  1. AniTalker: Animate Vivid and Diverse Talking Faces through Identity-Decoupled Facial Motion Encoding
    2024/05/06 by Tao Liu, Liu, Tao, Feilong Chen +11 · 21 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Face recognition and analysis
  2. VoiceFlow: Efficient Text-to-Speech with Rectified Flow Matching
    2023/09/10 by Guo, Yiwei, Du, Chenpeng, Ma, Ziyang +2 · 13 citations
    #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Human-Computer Interaction (cs.HC) #Sound (cs.SD) #electronic engineering #information engineering
  3. Language Model Can Listen While Speaking
    2024/08/05 by Ziyang Ma, Ma, Ziyang, Song, Yakun +12 · 17 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Human-Computer Interaction (cs.HC) #Natural Language Processing Techniques #Sound (cs.SD) #Speech and dialogue systems #electronic engineering #information engineering
  4. DiTAR: Diffusion Transformer Autoregressive Modeling for Speech Generation
    2025/02/06 by Jia, Dongya, Chen, Zhuo, Chen, Jiawei +8 · 23 citations
    #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Sound (cs.SD) #electronic engineering #information engineering
  5. Recent Advances in Discrete Speech Tokens: A Review
    2025/02/10 by Yiwei Guo, Guo, Yiwei, Zhihan Li +16 · 24 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Internet Traffic Analysis and Secure E-voting #Multimedia (cs.MM) #Signal Processing (eess.SP) #Sound (cs.SD) #Speech Recognition and Synthesis #electronic engineering #information engineering #semigroups and automata theory
  6. EmoDiff: Intensity Controllable Emotional Text-to-Speech with Soft-Label Guidance
    2022/11/17 by Yiwei Guo, Chenpeng Du, Guo, Yiwei +5 · 8 citations
    Computer Science · Psychology · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Human-Computer Interaction (cs.HC) #Machine Learning (cs.LG) #Phonetics and Phonology Research #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  7. Towards Universal Speech Discrete Tokens: A Case Study for ASR and TTS
    2023/09/14 by Yifan Yang, Yang, Yifan, Feiyu Shen +11 · 9 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and dialogue systems #Topic Modeling #electronic engineering #information engineering
  8. LSCodec: Low-Bitrate and Speaker-Decoupled Discrete Speech Codec
    2024/10/21 by Yiwei Guo, Zhihan Li, Guo, Yiwei +9 · 9 citations
    Computer Science · #Advanced Data Compression Techniques #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  9. DiffDub: Person-generic Visual Dubbing Using Inpainting Renderer with Diffusion Auto-encoder
    2023/11/03 by Tao Liu, Liu, Tao, Chenpeng Du +7 · 3 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Face recognition and analysis #Generative Adversarial Networks and Image Synthesis #Multimodal Machine Learning Applications
  10. Attention-Constrained Inference for Robust Decoder-Only Text-to-Speech
    2024/04/30 by Hankun Wang, Chenpeng Du, Wang, Hankun +9 · 2 citations
    Computer Science · #Speech Recognition and Synthesis
  11. GSTalker: Real-time Audio-Driven Talking Face Generation via Deformable Gaussian Splatting
    2024/04/29 by Chen, Bo, Hu, Shoukang, Chen, Qi +4 · 2 citations
    #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences
  12. Unsupervised word-level prosody tagging for controllable speech synthesis
    2022/02/15 by Guo, Yiwei, Du, Chenpeng, Yu, Kai · 1 citation
    #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Sound (cs.SD) #electronic engineering #information engineering
  13. DSE-TTS: Dual Speaker Embedding for Cross-Lingual Text-to-Speech
    2023/06/25 by Sen Liu, Yiwei Guo, Liu, Sen +7 · 1 citation
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Topic Modeling #electronic engineering #information engineering
  14. Acoustic BPE for Speech Generation with Discrete Tokens
    2023/10/23 by Feiyu Shen, Shen, Feiyu, Yiwei Guo +7 · 1 citation
    Computer Science · #Speech Recognition and Synthesis #Speech and dialogue systems #Topic Modeling
  15. MagiCodec: Simple Masked Gaussian-Injected Codec for High-Fidelity Reconstruction and Generation
    2025/05/31 by Song, Yakun, Chen, Jiawei, Zhuang, Xiaobin +9 · 3 citations
    #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Sound (cs.SD) #electronic engineering #information engineering
  16. vec2wav 2.0: Advancing Voice Conversion via Discrete Token Vocoders
    2024/09/03 by Yiwei Guo, Guo, Yiwei, Zhihan Li +13 · 1 citation
    Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  17. Towards Reliable Large Audio Language Model
    2025/05/25 by Ziyang Ma, Ma, Ziyang, Xiquan Li +17 · 2 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Human-Computer Interaction (cs.HC) #Multimedia (cs.MM) #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  18. Why Do Speech Language Models Fail to Generate Semantically Coherent Outputs? A Modality Evolving Perspective
    2024/12/22 by Wang, Hankun, Wang, Haoran, Guo, Yiwei +4 · 1 citation
    #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering