vix.ing · top · new · best · stats · spec

Xue, Liumeng

  1. Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens
    2025/03/03 by Xinsheng Wang, Mingqi Jiang, Ming Jiang +49 · 2 voices · 59 citations
    Computer Science · #Natural Language Processing Techniques #Speech Recognition and Synthesis #Speech and dialogue systems #cs.AI #cs.SD #eess.AS
  2. Llasa: Scaling Train-Time and Inference-Time Compute for Llama-based Speech Synthesis
    2025/02/06 by Ye, Zhen, Zhu, Xinfa, Chan, Chi-Min +17 · 38 citations
    #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Multimedia (cs.MM) #Sound (cs.SD) #electronic engineering #information engineering
  3. ChatMusician: Understanding and Generating Music Intrinsically with LLM
    2024/02/25 by Ruibin Yuan, Yuan, Ruibin, Y. F. Wang +66 · 16 citations
    Computer Science · Decision Sciences · Materials Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Machine Learning in Materials Science #Multimedia (cs.MM) #Scientific Computing and Data Management #Sound (cs.SD) #Topic Modeling #electronic engineering #information engineering
  4. Amphion: An Open-Source Audio, Music and Speech Generation Toolkit
    2023/12/15 by Xueyao Zhang, Liumeng Xue, Zhang, Xueyao +29 · 17 citations
    Computer Science · #Computational Physics and Python Applications #Speech Recognition and Synthesis #Music and Audio Processing
  5. WenetSpeech4TTS: A 12,800-hour Mandarin TTS Corpus for Large Speech Generation Model Benchmark
    2024/06/09 by Linhan Ma, Ma, Linhan, Dake Guo +16 · 15 citations
    Computer Science · Psychology · #Audio and Speech Processing (eess.AS) #FOS: Electrical engineering #Phonetics and Phonology Research #Speech Recognition and Synthesis #Speech and dialogue systems #electronic engineering #information engineering
  6. Single-Codec: Single-Codebook Speech Codec towards High-Performance Speech Generation
    2024/06/11 by Hanzhao Li, Li, Hanzhao, Liumeng Xue +15 · 14 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Electrical engineering #Speech Recognition and Synthesis #Speech and dialogue systems #electronic engineering #information engineering
  7. YuE: Scaling Open Foundation Models for Long-Form Music Generation
    2025/03/11 by Ruibin Yuan, Yuan, Ruibin, Shuyue Guo +110 · 23 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computer Graphics and Visualization Techniques #FOS: Computer and information sciences #FOS: Electrical engineering #Multimedia (cs.MM) #Music Technology and Sound Studies #Music and Audio Processing #Sound (cs.SD) #electronic engineering #information engineering
  8. Multi-Scale Sub-Band Constant-Q Transform Discriminator for High-Fidelity Vocoder
    2023/11/25 by Gu, Yicheng, Zhang, Xueyao, Xue, Liumeng +1 · 6 citations
    #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
  9. Expressive-VC: Highly Expressive Voice Conversion with Attention Fusion of Bottleneck and Perturbation Features
    2022/11/09 by Ning, Ziqian, Xie, Qicong, Zhu, Pengcheng +5 · 5 citations
    #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
  10. Controllable Emotion Transfer For End-to-End Speech Synthesis
    2020/11/17 by Tao Li, Li, Tao, Shan Yang +5 · 3 citations
    Computer Science · #Speech Recognition and Synthesis #Speech and Audio Processing #Advanced Data Compression Techniques
  11. AudioX: A Unified Framework for Anything-to-Audio Generation
    2025/03/13 by Zhaoyang Liu, Tian, Zeyue, Jin, Yizhu +10 · 11 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #FOS: Electrical engineering #Generative Adversarial Networks and Image Synthesis #Machine Learning (cs.LG) #Multimedia (cs.MM) #Music Technology and Sound Studies #Music and Audio Processing #Sound (cs.SD) #electronic engineering #information engineering
  12. Leveraging Diverse Semantic-based Audio Pretrained Models for Singing Voice Conversion
    2023/10/17 by Zhang, Xueyao, Fang, Zihao, Gu, Yicheng +5 · 4 citations
    #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
  13. SponTTS: modeling and transferring spontaneous style for TTS
    2023/11/13 by Li, Hanzhao, Zhu, Xinfa, Xue, Liumeng +3 · 2 citations
    #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
  14. Learning Noise-independent Speech Representation for High-quality Voice Conversion for Noisy Target Speakers
    2022/07/02 by Xue, Liumeng, Yang, Shan, Hu, Na +2 · 1 citation
    #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
  15. ParaTTS: Learning Linguistic and Prosodic Cross-sentence Information in Paragraph-based TTS
    2022/09/14 by Xue, Liumeng, Soong, Frank K., Zhang, Shaofei +1 · 1 citation
    #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
  16. HiGNN-TTS: Hierarchical Prosody Modeling with Graph Neural Networks for Expressive Long-form TTS
    2023/09/25 by Dake Guo, Guo, Dake, Xinfa Zhu +11 · 1 citation
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Topic Modeling #electronic engineering #information engineering
  17. An Investigation of Time-Frequency Representation Discriminators for High-Fidelity Vocoder
    2024/04/26 by Yicheng Gu, Xueyao Zhang, Gu, Yicheng +7 · 2 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sensor Technology and Measurement Systems #Signal Processing (eess.SP) #Sound (cs.SD) #electronic engineering #information engineering
  18. Text-aware and Context-aware Expressive Audiobook Speech Synthesis
    2024/06/09 by Guo, Dake, Zhu, Xinfa, Xue, Liumeng +3 · 1 citation
    #Audio and Speech Processing (eess.AS) #FOS: Electrical engineering #electronic engineering #information engineering