vix.ing · top · new · best · stats · spec

Ji, Shengpeng

  1. WavTokenizer: an Efficient Acoustic Discrete Codec Tokenizer for Audio Language Modeling
    2024/08/29 by Shengpeng Ji, Ziyue Karen Jiang, Ji, Shengpeng +30 · 41 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Multimedia (cs.MM) #Music and Audio Processing #Signal Processing (eess.SP) #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  2. FunAudioLLM: Voice Understanding and Generation Foundation Models for Natural Interaction Between Humans and LLMs
    2024/07/04 by An, Keyu, Chen, Qian, Deng, Chong +30 · 30 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #Speech and dialogue systems #electronic engineering #information engineering
  3. WavChat: A Survey of Spoken Dialogue Models
    2024/11/15 by Shengpeng Ji, Yifu Chen, Ji, Shengpeng +35 · 21 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Multi-Agent Systems and Negotiation #Multimedia (cs.MM) #Natural Language Processing Techniques #Sound (cs.SD) #Speech and dialogue systems #electronic engineering #information engineering
  4. MegaTTS 3: Sparse Alignment Enhanced Latent Diffusion Transformer for Zero-Shot Speech Synthesis
    2025/02/26 by Ziyue Karen Jiang, Yi Ren, Jiang, Ziyue +24 · 20 citations
    Computer Science · Medicine · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Face recognition and analysis #Machine Learning (cs.LG) #Sound (cs.SD) #Speech Recognition and Synthesis #Voice and Speech Disorders #electronic engineering #information engineering
  5. Mega-TTS: Zero-Shot Text-to-Speech at Scale with Intrinsic Inductive Bias
    2023/06/06 by Ziyue Karen Jiang, Jiang, Ziyue, Yi Ren +21 · 7 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  6. Mega-TTS 2: Boosting Prompting Mechanisms for Zero-Shot Speech Synthesis
    2023/07/14 by Ziyue Karen Jiang, Jiang, Ziyue, Jinglin Liu +21 · 7 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  7. ControlSpeech: Towards Simultaneous and Independent Zero-shot Speaker Cloning and Zero-shot Language Style Control
    2024/06/03 by Shengpeng Ji, Qian Chen, Ji, Shengpeng +18 · 8 citations
    Computer Science · #Speech Recognition and Synthesis #Natural Language Processing Techniques #Topic Modeling
  8. UniCodec: Unified Audio Codec with Single Domain-Adaptive Codebook
    2025/02/27 by Yidi Jiang, Qian Chen, Jiang, Yidi +15 · 10 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Generative Adversarial Networks and Image Synthesis #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  9. WavRAG: Audio-Integrated Retrieval Augmented Generation for Spoken Dialogue Models
    2025/02/20 by Y.W. Chen, Chen, Yifu, Shengpeng Ji +13 · 8 citations
    Computer Science · #Speech and dialogue systems #Topic Modeling #Natural Language Processing Techniques
  10. OmniSep: Unified Omni-Modality Sound Separation with Query-Mixup
    2024/10/28 by Cheng, Xize, Zheng, Siqi, Wang, Zehan +8 · 6 citations
    #Audio and Speech Processing (eess.AS) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #FOS: Electrical engineering #Multimedia (cs.MM) #Sound (cs.SD) #electronic engineering #information engineering
  11. LLaVA-MR: Large Language-and-Vision Assistant for Video Moment Retrieval
    2024/11/21 by Weisheng Lu, Jian Li, Lu, Weiheng +9 · 6 citations
    Computer Science · #Advanced Image and Video Retrieval Techniques #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Multimodal Machine Learning Applications #Video Analysis and Summarization
  12. Language-Codec: Bridging Discrete Codec Representations and Speech Language Models
    2024/02/19 by Ji, Shengpeng, Fang, Minghui, Zuo, Jialong +5 · 3 citations
    #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
  13. OmniChat: Enhancing Spoken Dialogue Systems with Scalable Synthetic Data for Diverse Scenarios
    2025/01/02 by Xize Cheng, Dongjie Fu, Cheng, Xize +25 · 6 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Human-Computer Interaction (cs.HC) #Multi-Agent Systems and Negotiation #Sound (cs.SD) #Speech and dialogue systems #Topic Modeling #electronic engineering #information engineering
  14. InSerter: Speech Instruction Following with Unsupervised Interleaved Pre-training
    2025/03/04 by Dingdong Wang, Wang, Dingdong, Jin Xu +14 · 6 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Human-Computer Interaction (cs.HC) #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Topic Modeling #electronic engineering #information engineering
  15. Enhancing Multimodal Unified Representations for Cross Modal Generalization
    2024/03/08 by Hai Huang, Yan Xia, Huang, Hai +15 · 3 citations
    Computer Science · #Natural Language Processing Techniques
  16. T2A-Feedback: Improving Basic Capabilities of Text-to-Audio Generation via Fine-grained AI Feedback
    2025/05/15 by Wang, Zehan, Lei, Ke, Zhu, Chen +8 · 3 citations
    #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
  17. Enhancing Expressive Voice Conversion with Discrete Pitch-Conditioned Flow Matching Model
    2025/02/08 by Jialong Zuo, Zuo, Jialong, Shengpeng Ji +19 · 3 citations
    Computer Science · #Speech Recognition and Synthesis #Speech and Audio Processing #Speech and dialogue systems
  18. MuVi: Video-to-Music Generation with Semantic Alignment and Rhythmic Synchronization
    2024/10/16 by Li, Ruiqi, Zheng, Siqi, Cheng, Xize +3 · 2 citations
    #Audio and Speech Processing (eess.AS) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #FOS: Electrical engineering #Multimedia (cs.MM) #Sound (cs.SD) #electronic engineering #information engineering
  19. CART: A Generative Cross-Modal Retrieval Framework with Coarse-To-Fine Semantic Modeling
    2024/06/25 by Fang, Minghui, Ji, Shengpeng, Zuo, Jialong +9 · 1 citation
    #FOS: Computer and information sciences #Information Retrieval (cs.IR)
  20. WavReward: Spoken Dialogue Models With Generalist Reward Evaluators
    2025/05/14 by Shengpeng Ji, Ji, Shengpeng, Yangzhuo Li +22 · 3 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Multi-Agent Systems and Negotiation #Multimedia (cs.MM) #Sound (cs.SD) #Speech and dialogue systems #Topic Modeling #electronic engineering #information engineering
  21. IRBridge: Solving Image Restoration Bridge with Pre-trained Generative Diffusion Models
    2025/05/30 by Wang, Hanting, Jin, Tao, Lin, Wang +4 · 3 citations
    #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences
  22. MobileSpeech: A Fast and High-Fidelity Framework for Mobile Zero-Shot Text-to-Speech
    2024/02/14 by Ji, Shengpeng, Jiang, Ziyue, Wang, Hanting +2 · 1 citation
    #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
  23. Astrea: A MOE-based Visual Understanding Model with Progressive Alignment
    2025/03/12 by Xiaoda Yang, Yang, Xiaoda, JunYu Lu +27 · 1 citation
    Computer Science · #Advanced Image and Video Retrieval Techniques #Artificial Intelligence (cs.AI) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Image Retrieval and Classification Techniques #Multimodal Machine Learning Applications