vix.ing · top · new · best · stats · spec

Zhizheng Wu

  1. NaturalSpeech 3: Zero-Shot Speech Synthesis with Factorized Codec and Diffusion Models
    2024/03/05 by Zeqian Ju, Ju, Zeqian, Yuancheng Wang +35 · 75 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  2. Emilia: An Extensive, Multilingual, and Diverse Speech Dataset for Large-Scale Speech Generation
    2024/07/07 by Haorui He, Zengqiang Shang, He, Haorui +25 · 86 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Speech Recognition and Synthesis #Speech and dialogue systems #Topic Modeling #electronic engineering #information engineering
  3. FoleyCrafter: Bring Silent Videos to Life with Lifelike and Synchronized Sounds
    2024/07/01 by Yiming Zhang, Zhang, Yiming, Yicheng Gu +11 · 1 voice · 27 citations
    Neuroscience · Computer Science · #cs.CV #cs.SD #eess.AS
  4. MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer
    2024/09/01 by Yuancheng Wang, Haoyue Zhan, Wang, Yuancheng +16 · 63 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #electronic engineering #information engineering
  5. SD-Eval: A Benchmark Dataset for Spoken Dialogue Understanding Beyond Words
    2024/06/19 by Junyi Ao, Ao, Junyi, Yuancheng Wang +15 · 18 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Sound (cs.SD) #Speech and dialogue systems #Topic Modeling #electronic engineering #information engineering
  6. Amphion: An Open-Source Audio, Music and Speech Generation Toolkit
    2023/12/15 by Xueyao Zhang, Zhang, Xueyao, Liumeng Xue +29 · 17 citations
    Computer Science · #Computational Physics and Python Applications #Speech Recognition and Synthesis #Music and Audio Processing
  7. Investigating gated recurrent neural networks for speech synthesis
    2016/01/11 by Zhizheng Wu, Wu, Zhizheng, Simon King +1 · 1 voice
    #cs.CL #cs.NE
  8. AnyEnhance: A Unified Generative Model with Prompt-Guidance and Self-Critic for Voice Enhancement
    2025/01/26 by J. S. Zhang, Jing Yang, Zhang, Junan +12 · 19 citations
    Computer Science · #Speech and dialogue systems #Speech Recognition and Synthesis #Speech and Audio Processing
  9. Foundation Models for Music: A Survey
    2024/08/26 by Yinghao Ma, Ma, Yinghao, Anders Øland +81 · 9 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Music Technology and Sound Studies #Sound (cs.SD) #electronic engineering #information engineering
  10. AudioTime: A Temporally-aligned Audio-text Benchmark Dataset
    2024/07/03 by Zeyu Xie, Xuenan Xu, Xie, Zeyu +5 · 8 citations
    Computer Science · #Music and Audio Processing #Speech Recognition and Synthesis
  11. Spoofing and countermeasures for speaker verification: A survey
    2015/02/01 by Zhizheng Wu, Nicholas Evans, Tomi Kinnunen +3 · 2 citations
  12. Accented Text-to-Speech Synthesis with Limited Data
    2023/05/08 by Xuehao Zhou, Zhou, Xuehao, Mingyang Zhang +7 · 4 citations
    Computer Science · Psychology · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Phonetics and Phonology Research #Sound (cs.SD) #Speech Recognition and Synthesis #electronic engineering #information engineering
  13. Metis: A Foundation Speech Generation Model with Masked Generative Pre-training
    2025/02/05 by Yuancheng Wang, Wang, Yuancheng, Jiachen Zheng +9 · 11 citations
    Computer Science · #Speech and dialogue systems #Speech Recognition and Synthesis #Natural Language Processing Techniques
  14. Emilia: A Large-Scale, Extensive, Multilingual, and Diverse Dataset for Speech Generation
    2025/01/27 by Haorui He, He, Haorui, Zengqiang Shang +25 · 9 citations
    Computer Science · #Speech Recognition and Synthesis #Speech and dialogue systems
  15. PicoAudio: Enabling Precise Timestamp and Frequency Controllability of Audio Events in Text-to-audio Generation
    2024/07/03 by Zeyu Xie, Xie, Zeyu, Xuenan Xu +5 · 2 citations
    Computer Science · #68Txx #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #I.2 #Music Technology and Sound Studies #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #electronic engineering #information engineering
  16. SingNet: Towards a Large-Scale, Diverse, and In-the-Wild Singing Voice Dataset
    2025/05/14 by Yicheng Gu, Gu, Yicheng, Chaoren Wang +11 · 6 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  17. Investigating Neural Audio Codecs for Speech Language Model-Based Speech Generation
    2024/09/06 by Jiaqi Li, Li, Jiaqi, Dongmei Wang +29 · 3 citations
    Computer Science · #Speech Recognition and Synthesis #Music and Audio Processing
  18. Overview of the Amphion Toolkit (v0.2)
    2025/01/26 by Jiaqi Li, Li, Jiaqi, Xueyao Zhang +20 · 5 citations
    Physics and Astronomy · Engineering · #Particle physics theoretical and experimental studies #Quantum Chromodynamics and Particle Interactions #Superconducting Materials and Applications
  19. Vevo2: A Unified and Controllable Framework for Speech and Singing Voice Generation
    2025/08/22 by Xueyao Zhang, Zhang, Xueyao, J. S. Zhang +13 · 4 citations
    Computer Science · Medicine · Psychology · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Phonetics and Phonology Research #Sound (cs.SD) #Speech Recognition and Synthesis #Voice and Speech Disorders
  20. An Investigation of Time-Frequency Representation Discriminators for High-Fidelity Vocoder
    2024/04/26 by Yicheng Gu, Gu, Yicheng, Xueyao Zhang +7 · 2 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sensor Technology and Measurement Systems #Signal Processing (eess.SP) #Sound (cs.SD) #electronic engineering #information engineering
  21. CoAVT: A Cognition-Inspired Unified Audio-Visual-Text Pre-Training Model for Multimodal Processing
    2024/01/22 by Xianghu Yue, Xiaohai Tian, Yue, Xianghu +8 · 1 citation
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Image and Video Processing (eess.IV) #Multimedia (cs.MM) #Multimodal Machine Learning Applications #Music and Audio Processing #Sound (cs.SD) #Video Analysis and Summarization #electronic engineering #information engineering
  22. Multi-Scale Accent Modeling and Disentangling for Multi-Speaker Multi-Accent Text-to-Speech Synthesis
    2024/06/16 by Xuehao Zhou, Mingyang Zhang, Zhou, Xuehao +7 · 1 citation
    Computer Science · #Speech and Audio Processing #Speech Recognition and Synthesis
  23. SpMis: An Investigation of Synthetic Spoken Misinformation Detection
    2024/09/17 by Peizhuo Liu, Liu, Peizhuo, Li Wang +15 · 1 citation
    Computer Science · Social Sciences · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Hate Speech and Cyberbullying Detection #Misinformation and Its Impacts #Sentiment Analysis and Opinion Mining
  24. Neurodyne: Neural Pitch Manipulation with Representation Learning and Cycle-Consistency GAN
    2025/05/21 by Yicheng Gu, Gu, Yicheng, Chaoren Wang +5 · 2 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music Technology and Sound Studies #Music and Audio Processing #Sound (cs.SD) #Speech and Audio Processing #electronic engineering #information engineering
  25. Is One Score Enough? Assessing Singing Quality of Songs with Temporal Score Curves
    2026/07/18 by Yishan Lv, Jing Luo, Xinyu Yang +1
    #cs.SD #cs.MM #eess.AS
  26. Anysynth:Zero-Shot Instrument Cloning via In-Context Learning and Asymmetric Hierarchical Guidance
    2026/07/13 by Chong Jing, Junan Zhang, Jing Yang +3
    #cs.SD
  27. Teffic-Audio: Tell Fact from Fiction
    2026/07/30 by Wan Lin, Li Wang, Jindong Wang +2
    Computer Science · #cs.SD #cs.AI
  28. SimulS2ST-Omni: Data-Efficient Streaming Speech-to-Speech Translation via Explicit Trajectory Supervision
    2026/07/22 by Rongshen He, Xinyu Liang, Dekun Chen +3
    #cs.SD