Xinfa Zhu
- Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens
2025/03/03 by Xinsheng Wang, Wang, Xinsheng, Mingqi Jiang +49 · 2 voices · 59 citations
Computer Science · #Natural Language Processing Techniques #Speech Recognition and Synthesis #Speech and dialogue systems #cs.AI #cs.SD #eess.AS
- Qwen3-Omni Technical Report
2025/09/22 by Xu Jin, Jin Xu, Xu, Jin +78 · 1 voice · 66 citations
Computer Science · Engineering · #Generative Adversarial Networks and Image Synthesis #Multimodal Machine Learning Applications #Speech and Audio Processing #cs.AI #cs.CL #cs.CV #eess.AS
- SELM: Speech Enhancement Using Discrete Tokens and Language Models
2023/12/15 by Ziqian Wang, Wang, Ziqian, Xinfa Zhu +11 · 16 citations
Computer Science · Medicine · #Speech and Audio Processing #Speech Recognition and Synthesis #Voice and Speech Disorders
- Single-Codec: Single-Codebook Speech Codec towards High-Performance Speech Generation
2024/06/11 by Hanzhao Li, Liumeng Xue, Li, Hanzhao +15 · 14 citations
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Electrical engineering #Speech Recognition and Synthesis #Speech and dialogue systems #electronic engineering #information engineering
- METTS: Multilingual Emotional Text-to-Speech by Cross-speaker and Cross-lingual Emotion Transfer
2023/07/29 by Xinfa Zhu, Yi Lei, Zhu, Xinfa +11 · 6 citations
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Electrical engineering #Speech Recognition and Synthesis #electronic engineering #information engineering
- FlowSE: Efficient and High-Quality Speech Enhancement via Flow Matching
2025/05/26 by Ziqian Wang, Wang, Ziqian, Liu, Zikai +14 · 11 citations
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Electrical engineering #Music and Audio Processing #Signal Processing (eess.SP) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- DiCLET-TTS: Diffusion Model based Cross-lingual Emotion Transfer for Text-to-Speech -- A Study between English and Mandarin
2023/09/02 by Tao Li, Li, Tao, Chenxu Hu +13 · 3 citations
Psychology · Computer Science · #Phonetics and Phonology Research #Speech Recognition and Synthesis #Sentiment Analysis and Opinion Mining
- KALL-E:Autoregressive Speech Synthesis with Next-Distribution Prediction
2024/12/22 by Xinfa Zhu, Xia, Kangxiang, Zhu, Xinfa +6 · 5 citations
Computer Science · Psychology · #Speech Recognition and Synthesis #Speech and Audio Processing #Phonetics and Phonology Research
- Vec-Tok-VC+: Residual-enhanced Robust Zero-shot Voice Conversion with Progressive Constraints in a Dual-mode Training Strategy
2024/06/14 by Linhan Ma, Ma, Linhan, Xinfa Zhu +13 · 3 citations
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #Speech and dialogue systems #electronic engineering #information engineering
- Zero-Shot Emotion Transfer For Cross-Lingual Speech Synthesis
2023/10/06 by Yuke Li, Li, Yuke, Xinfa Zhu +11 · 2 citations
Computer Science · #Speech Recognition and Synthesis #Speech and Audio Processing
- Multi-Speaker Expressive Speech Synthesis via Multiple Factors Decoupling
2022/11/19 by Xinfa Zhu, Yi Lei, Zhu, Xinfa +9 · 1 citation
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- HiGNN-TTS: Hierarchical Prosody Modeling with Graph Neural Networks for Expressive Long-form TTS
2023/09/25 by Dake Guo, Xinfa Zhu, Guo, Dake +11 · 1 citation
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Topic Modeling #electronic engineering #information engineering
- Accent-VITS:accent transfer for end-to-end TTS
2023/12/28 by Linhan Ma, Ma, Linhan, Yongmao Zhang +11 · 1 citation
Computer Science · Psychology · Medicine · #Speech Recognition and Synthesis #Phonetics and Phonology Research #Voice and Speech Disorders
- Qwen-Music Technical Report
2026/07/27 by Jin Xu, Kangdi Wang, Ruibin Yuan +24
#cs.SD