Sitong Cheng
- Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens
2025/03/03 by Xinsheng Wang, Mingqi Jiang, Ming Jiang +49 · 2 voices · 65 citations
Computer Science · #Natural Language Processing Techniques #Speech Recognition and Synthesis #Speech and dialogue systems #cs.AI #cs.SD #eess.AS
- CN-CELEB: a challenging Chinese speaker recognition dataset
2019/10/31 by Yue Fan, Fan, Yue, Jiawen Kang +17 · 11 citations
Computer Science · #Speech Recognition and Synthesis #Music and Audio Processing #Speech and Audio Processing
- Both Ears Wide Open: Towards Language-Driven Spatial Audio Generation
2024/10/14 by Peiwen Sun, Sun, Peiwen, Sitong Cheng +12 · 14 citations
Computer Science · #Audio and Speech Processing (eess.AS) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #FOS: Electrical engineering #Music Technology and Sound Studies #Music and Audio Processing #Sound (cs.SD) #electronic engineering #information engineering