Shengpeng Ji
- WavTokenizer: an Efficient Acoustic Discrete Codec Tokenizer for Audio Language Modeling
2024/08/29 by Shengpeng Ji, Ziyue Karen Jiang, Ji, Shengpeng +33 · 1 voice · 48 citations
Computer Science · Engineering · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Multimedia (cs.MM) #Music and Audio Processing #Signal Processing (eess.SP) #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #cs.LG #cs.MM #cs.SD #eess.AS #eess.SP #electronic engineering #information engineering
- WavChat: A Survey of Spoken Dialogue Models
2024/11/15 by Shengpeng Ji, Ji, Shengpeng, Yifu Chen +35 · 28 citations
Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Multi-Agent Systems and Negotiation #Multimedia (cs.MM) #Natural Language Processing Techniques #Sound (cs.SD) #Speech and dialogue systems #electronic engineering #information engineering
- MegaTTS 3: Sparse Alignment Enhanced Latent Diffusion Transformer for Zero-Shot Speech Synthesis
2025/02/26 by Ziyue Karen Jiang, Jiang, Ziyue, Yi Ren +24 · 23 citations
Computer Science · Medicine · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Face recognition and analysis #Machine Learning (cs.LG) #Sound (cs.SD) #Speech Recognition and Synthesis #Voice and Speech Disorders #electronic engineering #information engineering
- Mega-TTS: Zero-Shot Text-to-Speech at Scale with Intrinsic Inductive Bias
2023/06/06 by Ziyue Karen Jiang, Yi Ren, Jiang, Ziyue +21 · 7 citations
Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- ControlSpeech: Towards Simultaneous and Independent Zero-shot Speaker Cloning and Zero-shot Language Style Control
2024/06/03 by Shengpeng Ji, Qian Chen, Ji, Shengpeng +18 · 8 citations
Computer Science · #Speech Recognition and Synthesis #Natural Language Processing Techniques #Topic Modeling
- UniCodec: Unified Audio Codec with Single Domain-Adaptive Codebook
2025/02/27 by Yidi Jiang, Jiang, Yidi, Qian Chen +15 · 12 citations
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Generative Adversarial Networks and Image Synthesis #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- WavRAG: Audio-Integrated Retrieval Augmented Generation for Spoken Dialogue Models
2025/02/20 by Y.W. Chen, Shengpeng Ji, Chen, Yifu +13 · 8 citations
Computer Science · #Speech and dialogue systems #Topic Modeling #Natural Language Processing Techniques
- LLaVA-MR: Large Language-and-Vision Assistant for Video Moment Retrieval
2024/11/21 by Weisheng Lu, Lu, Weiheng, Jian Li +9 · 7 citations
Computer Science · #Advanced Image and Video Retrieval Techniques #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Multimodal Machine Learning Applications #Video Analysis and Summarization
- Language-Codec: Bridging Discrete Codec Representations and Speech Language Models
2024/02/19 by Shengpeng Ji, Ji, Shengpeng, Minghui Fang +9 · 4 citations
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and dialogue systems #electronic engineering #information engineering
- InSerter: Speech Instruction Following with Unsupervised Interleaved Pre-training
2025/03/04 by Dingdong Wang, Jin Xu, Wang, Dingdong +14 · 7 citations
Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Human-Computer Interaction (cs.HC) #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Topic Modeling #electronic engineering #information engineering
- OmniChat: Enhancing Spoken Dialogue Systems with Scalable Synthetic Data for Diverse Scenarios
2025/01/02 by Xize Cheng, Cheng, Xize, Dongjie Fu +25 · 6 citations
Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Human-Computer Interaction (cs.HC) #Multi-Agent Systems and Negotiation #Sound (cs.SD) #Speech and dialogue systems #Topic Modeling #electronic engineering #information engineering
- Enhancing Multimodal Unified Representations for Cross Modal Generalization
2024/03/08 by Hai Huang, Yan Xia, Huang, Hai +15 · 4 citations
Computer Science · #Natural Language Processing Techniques
- Enhancing Expressive Voice Conversion with Discrete Pitch-Conditioned Flow Matching Model
2025/02/08 by Jialong Zuo, Zuo, Jialong, Shengpeng Ji +19 · 4 citations
Computer Science · #Speech Recognition and Synthesis #Speech and Audio Processing #Speech and dialogue systems
- WavReward: Spoken Dialogue Models With Generalist Reward Evaluators
2025/05/14 by Shengpeng Ji, Ji, Shengpeng, Yangzhuo Li +22 · 3 citations
Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Multi-Agent Systems and Negotiation #Multimedia (cs.MM) #Sound (cs.SD) #Speech and dialogue systems #Topic Modeling #electronic engineering #information engineering
- IRBridge: Solving Image Restoration Bridge with Pre-trained Generative Diffusion Models
2025/05/30 by Hanting Wang, Wang, Hanting, Tao Jin +10 · 5 citations
Computer Science · #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Generative Adversarial Networks and Image Synthesis #Image and Signal Denoising Methods
- X3-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment
2026/07/23 by Dongjie Fu, Di Cao, Xize Cheng +6
#cs.LG