Zhizheng Wu
- NaturalSpeech 3: Zero-Shot Speech Synthesis with Factorized Codec and Diffusion Models
2024/03/05 by Zeqian Ju, Ju, Zeqian, Yuancheng Wang +35 · 75 citations
Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- Emilia: An Extensive, Multilingual, and Diverse Speech Dataset for Large-Scale Speech Generation
2024/07/07 by Haorui He, Zengqiang Shang, He, Haorui +25 · 86 citations
Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Speech Recognition and Synthesis #Speech and dialogue systems #Topic Modeling #electronic engineering #information engineering
- FoleyCrafter: Bring Silent Videos to Life with Lifelike and Synchronized Sounds
2024/07/01 by Yiming Zhang, Zhang, Yiming, Yicheng Gu +11 · 1 voice · 27 citations
Neuroscience · Computer Science · #cs.CV #cs.SD #eess.AS
- MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer
2024/09/01 by Yuancheng Wang, Haoyue Zhan, Wang, Yuancheng +16 · 63 citations
Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #electronic engineering #information engineering
- SD-Eval: A Benchmark Dataset for Spoken Dialogue Understanding Beyond Words
2024/06/19 by Junyi Ao, Ao, Junyi, Yuancheng Wang +15 · 18 citations
Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Sound (cs.SD) #Speech and dialogue systems #Topic Modeling #electronic engineering #information engineering
- Amphion: An Open-Source Audio, Music and Speech Generation Toolkit
2023/12/15 by Xueyao Zhang, Zhang, Xueyao, Liumeng Xue +29 · 17 citations
Computer Science · #Computational Physics and Python Applications #Speech Recognition and Synthesis #Music and Audio Processing
- Investigating gated recurrent neural networks for speech synthesis
2016/01/11 by Zhizheng Wu, Wu, Zhizheng, Simon King +1 · 1 voice
#cs.CL #cs.NE
- AnyEnhance: A Unified Generative Model with Prompt-Guidance and Self-Critic for Voice Enhancement
2025/01/26 by J. S. Zhang, Jing Yang, Zhang, Junan +12 · 19 citations
Computer Science · #Speech and dialogue systems #Speech Recognition and Synthesis #Speech and Audio Processing
- Foundation Models for Music: A Survey
2024/08/26 by Yinghao Ma, Ma, Yinghao, Anders Øland +81 · 9 citations
Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Music Technology and Sound Studies #Sound (cs.SD) #electronic engineering #information engineering
- AudioTime: A Temporally-aligned Audio-text Benchmark Dataset
2024/07/03 by Zeyu Xie, Xuenan Xu, Xie, Zeyu +5 · 8 citations
Computer Science · #Music and Audio Processing #Speech Recognition and Synthesis
- Spoofing and countermeasures for speaker verification: A survey
2015/02/01 by Zhizheng Wu, Nicholas Evans, Tomi Kinnunen +3 · 2 citations
- Accented Text-to-Speech Synthesis with Limited Data
2023/05/08 by Xuehao Zhou, Zhou, Xuehao, Mingyang Zhang +7 · 4 citations
Computer Science · Psychology · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Phonetics and Phonology Research #Sound (cs.SD) #Speech Recognition and Synthesis #electronic engineering #information engineering
- Metis: A Foundation Speech Generation Model with Masked Generative Pre-training
2025/02/05 by Yuancheng Wang, Wang, Yuancheng, Jiachen Zheng +9 · 11 citations
Computer Science · #Speech and dialogue systems #Speech Recognition and Synthesis #Natural Language Processing Techniques
- Emilia: A Large-Scale, Extensive, Multilingual, and Diverse Dataset for Speech Generation
2025/01/27 by Haorui He, He, Haorui, Zengqiang Shang +25 · 9 citations
Computer Science · #Speech Recognition and Synthesis #Speech and dialogue systems
- PicoAudio: Enabling Precise Timestamp and Frequency Controllability of Audio Events in Text-to-audio Generation
2024/07/03 by Zeyu Xie, Xie, Zeyu, Xuenan Xu +5 · 2 citations
Computer Science · #68Txx #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #I.2 #Music Technology and Sound Studies #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #electronic engineering #information engineering
- SingNet: Towards a Large-Scale, Diverse, and In-the-Wild Singing Voice Dataset
2025/05/14 by Yicheng Gu, Gu, Yicheng, Chaoren Wang +11 · 6 citations
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- Investigating Neural Audio Codecs for Speech Language Model-Based Speech Generation
2024/09/06 by Jiaqi Li, Li, Jiaqi, Dongmei Wang +29 · 3 citations
Computer Science · #Speech Recognition and Synthesis #Music and Audio Processing
- Overview of the Amphion Toolkit (v0.2)
2025/01/26 by Jiaqi Li, Li, Jiaqi, Xueyao Zhang +20 · 5 citations
Physics and Astronomy · Engineering · #Particle physics theoretical and experimental studies #Quantum Chromodynamics and Particle Interactions #Superconducting Materials and Applications
- Vevo2: A Unified and Controllable Framework for Speech and Singing Voice Generation
2025/08/22 by Xueyao Zhang, Zhang, Xueyao, J. S. Zhang +13 · 4 citations
Computer Science · Medicine · Psychology · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Phonetics and Phonology Research #Sound (cs.SD) #Speech Recognition and Synthesis #Voice and Speech Disorders
- An Investigation of Time-Frequency Representation Discriminators for High-Fidelity Vocoder
2024/04/26 by Yicheng Gu, Gu, Yicheng, Xueyao Zhang +7 · 2 citations
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sensor Technology and Measurement Systems #Signal Processing (eess.SP) #Sound (cs.SD) #electronic engineering #information engineering
- CoAVT: A Cognition-Inspired Unified Audio-Visual-Text Pre-Training Model for Multimodal Processing
2024/01/22 by Xianghu Yue, Xiaohai Tian, Yue, Xianghu +8 · 1 citation
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Image and Video Processing (eess.IV) #Multimedia (cs.MM) #Multimodal Machine Learning Applications #Music and Audio Processing #Sound (cs.SD) #Video Analysis and Summarization #electronic engineering #information engineering
- Multi-Scale Accent Modeling and Disentangling for Multi-Speaker Multi-Accent Text-to-Speech Synthesis
2024/06/16 by Xuehao Zhou, Mingyang Zhang, Zhou, Xuehao +7 · 1 citation
Computer Science · #Speech and Audio Processing #Speech Recognition and Synthesis
- SpMis: An Investigation of Synthetic Spoken Misinformation Detection
2024/09/17 by Peizhuo Liu, Liu, Peizhuo, Li Wang +15 · 1 citation
Computer Science · Social Sciences · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Hate Speech and Cyberbullying Detection #Misinformation and Its Impacts #Sentiment Analysis and Opinion Mining
- Neurodyne: Neural Pitch Manipulation with Representation Learning and Cycle-Consistency GAN
2025/05/21 by Yicheng Gu, Gu, Yicheng, Chaoren Wang +5 · 2 citations
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music Technology and Sound Studies #Music and Audio Processing #Sound (cs.SD) #Speech and Audio Processing #electronic engineering #information engineering
- Is One Score Enough? Assessing Singing Quality of Songs with Temporal Score Curves
2026/07/18 by Yishan Lv, Jing Luo, Xinyu Yang +1
#cs.SD #cs.MM #eess.AS
- Anysynth:Zero-Shot Instrument Cloning via In-Context Learning and Asymmetric Hierarchical Guidance
2026/07/13 by Chong Jing, Junan Zhang, Jing Yang +3
#cs.SD
- Teffic-Audio: Tell Fact from Fiction
2026/07/30 by Wan Lin, Li Wang, Jindong Wang +2
Computer Science · #cs.SD #cs.AI
- SimulS2ST-Omni: Data-Efficient Streaming Speech-to-Speech Translation via Explicit Trajectory Supervision
2026/07/22 by Rongshen He, Xinyu Liang, Dekun Chen +3
#cs.SD