vix.ing · top · new · best · stats · spec

Xinfa Zhu

  1. Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens
    2025/03/03 by Xinsheng Wang, Wang, Xinsheng, Mingqi Jiang +49 · 2 voices · 59 citations
    Computer Science · #Natural Language Processing Techniques #Speech Recognition and Synthesis #Speech and dialogue systems #cs.AI #cs.SD #eess.AS
  2. Qwen3-Omni Technical Report
    2025/09/22 by Xu Jin, Jin Xu, Xu, Jin +78 · 1 voice · 66 citations
    Computer Science · Engineering · #Generative Adversarial Networks and Image Synthesis #Multimodal Machine Learning Applications #Speech and Audio Processing #cs.AI #cs.CL #cs.CV #eess.AS
  3. SELM: Speech Enhancement Using Discrete Tokens and Language Models
    2023/12/15 by Ziqian Wang, Wang, Ziqian, Xinfa Zhu +11 · 16 citations
    Computer Science · Medicine · #Speech and Audio Processing #Speech Recognition and Synthesis #Voice and Speech Disorders
  4. Single-Codec: Single-Codebook Speech Codec towards High-Performance Speech Generation
    2024/06/11 by Hanzhao Li, Liumeng Xue, Li, Hanzhao +15 · 14 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Electrical engineering #Speech Recognition and Synthesis #Speech and dialogue systems #electronic engineering #information engineering
  5. METTS: Multilingual Emotional Text-to-Speech by Cross-speaker and Cross-lingual Emotion Transfer
    2023/07/29 by Xinfa Zhu, Yi Lei, Zhu, Xinfa +11 · 6 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Electrical engineering #Speech Recognition and Synthesis #electronic engineering #information engineering
  6. FlowSE: Efficient and High-Quality Speech Enhancement via Flow Matching
    2025/05/26 by Ziqian Wang, Wang, Ziqian, Liu, Zikai +14 · 11 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Electrical engineering #Music and Audio Processing #Signal Processing (eess.SP) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  7. DiCLET-TTS: Diffusion Model based Cross-lingual Emotion Transfer for Text-to-Speech -- A Study between English and Mandarin
    2023/09/02 by Tao Li, Li, Tao, Chenxu Hu +13 · 3 citations
    Psychology · Computer Science · #Phonetics and Phonology Research #Speech Recognition and Synthesis #Sentiment Analysis and Opinion Mining
  8. KALL-E:Autoregressive Speech Synthesis with Next-Distribution Prediction
    2024/12/22 by Xinfa Zhu, Xia, Kangxiang, Zhu, Xinfa +6 · 5 citations
    Computer Science · Psychology · #Speech Recognition and Synthesis #Speech and Audio Processing #Phonetics and Phonology Research
  9. Vec-Tok-VC+: Residual-enhanced Robust Zero-shot Voice Conversion with Progressive Constraints in a Dual-mode Training Strategy
    2024/06/14 by Linhan Ma, Ma, Linhan, Xinfa Zhu +13 · 3 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #Speech and dialogue systems #electronic engineering #information engineering
  10. Zero-Shot Emotion Transfer For Cross-Lingual Speech Synthesis
    2023/10/06 by Yuke Li, Li, Yuke, Xinfa Zhu +11 · 2 citations
    Computer Science · #Speech Recognition and Synthesis #Speech and Audio Processing
  11. Multi-Speaker Expressive Speech Synthesis via Multiple Factors Decoupling
    2022/11/19 by Xinfa Zhu, Yi Lei, Zhu, Xinfa +9 · 1 citation
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  12. HiGNN-TTS: Hierarchical Prosody Modeling with Graph Neural Networks for Expressive Long-form TTS
    2023/09/25 by Dake Guo, Xinfa Zhu, Guo, Dake +11 · 1 citation
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Topic Modeling #electronic engineering #information engineering
  13. Accent-VITS:accent transfer for end-to-end TTS
    2023/12/28 by Linhan Ma, Ma, Linhan, Yongmao Zhang +11 · 1 citation
    Computer Science · Psychology · Medicine · #Speech Recognition and Synthesis #Phonetics and Phonology Research #Voice and Speech Disorders
  14. Qwen-Music Technical Report
    2026/07/27 by Jin Xu, Kangdi Wang, Ruibin Yuan +24
    #cs.SD