Dongchao Yang
- AudioGPT: Understanding and Generating Speech, Music, Sound, and Talking Head
2023/04/25 by Rongjie Huang, Huang, Rongjie, Mingze Li +23 · 1 voice · 52 citations
Computer Science · #Music and Audio Processing #Speech Recognition and Synthesis #Topic Modeling
- NaturalSpeech 3: Zero-Shot Speech Synthesis with Factorized Codec and Diffusion Models
2024/03/05 by Zeqian Ju, Ju, Zeqian, Yuancheng Wang +35 · 75 citations
Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- Make-An-Audio: Text-To-Audio Generation with Prompt-Enhanced Diffusion Models
2023/01/30 by Rongjie Huang, Jiawei Huang, Huang, Rongjie +17 · 42 citations
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Generative Adversarial Networks and Image Synthesis #Machine Learning (cs.LG) #Multimedia (cs.MM) #Music Technology and Sound Studies #Music and Audio Processing #Sound (cs.SD) #electronic engineering #information engineering
- Diffsound: Discrete Diffusion Model for Text-to-sound Generation
2022/07/20 by Dongchao Yang, Yang, Dongchao, Jianwei Yu +11 · 26 citations
Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music Technology and Sound Studies #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #electronic engineering #information engineering
- HiFi-Codec: Group-residual Vector quantization for High Fidelity Audio Codec
2023/05/04 by Dongchao Yang, Yang, Dongchao, Songxiang Liu +9 · 31 citations
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- Make-An-Audio 2: Temporal-Enhanced Text-to-Audio Generation
2023/05/29 by Jiawei Huang, Huang, Jiawei, Yi Ren +17 · 21 citations
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Multimedia (cs.MM) #Music Technology and Sound Studies #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #electronic engineering #information engineering
- InstructTTS: Modelling Expressive TTS in Discrete Latent Space with Natural Language Style Prompt
2023/01/31 by Dongchao Yang, Yang, Dongchao, Songxiang Liu +7 · 14 citations
Computer Science · #Speech Recognition and Synthesis #Speech and Audio Processing #Music and Audio Processing
- MMSU: A Massive Multi-task Spoken Language Understanding and Reasoning Benchmark
2025/06/05 by Dingdong Wang, Wang, Dingdong, Junan Li +11 · 31 citations
Computer Science · Psychology · #Speech Recognition and Synthesis #Emotion and Mood Recognition #Speech and dialogue systems
- UniAudio 1.5: Large Language Model-driven Audio Codec is A Few-shot Audio Task Learner
2024/06/14 by Dongchao Yang, Yang, Dongchao, Haohan Guo +13 · 10 citations
Computer Science · #Music and Audio Processing #Speech Recognition and Synthesis #Speech and Audio Processing
- PromptTTS 2: Describing and Generating Voices with Text Prompt
2023/09/05 by Yichong Leng, Zhifang Guo, Leng, Yichong +27 · 7 citations
Computer Science · #Speech Recognition and Synthesis #Topic Modeling #Natural Language Processing Techniques
- SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline
2025/05/25 by Helin Wang, Jiarui Hai, Wang, Helin +17 · 1 voice · 4 citations
Computer Science · Engineering · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #cs.AI #cs.SD #eess.AS #electronic engineering #information engineering
- Make-A-Voice: Unified Voice Synthesis With Discrete Representation
2023/05/30 by Rongjie Huang, Huang, Rongjie, Chunlei Zhang +17 · 5 citations
Computer Science · #Speech Recognition and Synthesis #Music and Audio Processing #Topic Modeling
- RALL-E: Robust Codec Language Modeling with Chain-of-Thought Prompting for Text-to-Speech Synthesis
2024/04/04 by Detai Xin, Xu Tan, Xin, Detai +19 · 4 citations
Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Topic Modeling #electronic engineering #information engineering
- SimpleSpeech 2: Towards Simple and Efficient Text-to-Speech with Flow-based Scalar Latent Transformer Diffusion Models
2024/08/25 by Dongchao Yang, Rongjie Huang, Yang, Dongchao +13 · 5 citations
Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Topic Modeling #electronic engineering #information engineering
- InSerter: Speech Instruction Following with Unsupervised Interleaved Pre-training
2025/03/04 by Dingdong Wang, Wang, Dingdong, Jin Xu +14 · 7 citations
Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Human-Computer Interaction (cs.HC) #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Topic Modeling #electronic engineering #information engineering
- Improving Target Sound Extraction with Timestamp Information
2022/04/02 by Helin Wang, Dongchao Yang, Wang, Helin +7 · 2 citations
Computer Science · Engineering · #Acoustic Wave Phenomena Research #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech and Audio Processing #electronic engineering #information engineering
- AudioComposer: Towards Fine-grained Audio Generation with Natural Language Descriptions
2024/09/19 by Yuanyuan Wang, Wang, Yuanyuan, Hangting Chen +7 · 4 citations
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music Technology and Sound Studies #Music and Audio Processing #Sound (cs.SD) #Speech and Audio Processing #electronic engineering #information engineering
- ALMTokenizer: A Low-bitrate and Semantic-rich Audio Codec Tokenizer for Audio Language Modeling
2025/04/14 by Dongchao Yang, Songxiang Liu, Yang, Dongchao +21 · 8 citations
Computer Science · #FOS: Computer and information sciences #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing
- VARGPT-v1.1: Improve Visual Autoregressive Large Unified Model via Iterative Instruction Tuning and Reinforcement Learning
2025/04/03 by Xianwei Zhuang, Yuxin Xie, Zhuang, Xianwei +13 · 7 citations
Computer Science · Biochemistry, Genetics and Molecular Biology · #Generative Adversarial Networks and Image Synthesis #Multimodal Machine Learning Applications #Cell Image Analysis Techniques
- SoCodec: A Semantic-Ordered Multi-Stream Speech Codec for Efficient Language Model Based Text-to-Speech Synthesis
2024/09/02 by Haohan Guo, Guo, Haohan, Fenglong Xie +11 · 3 citations
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and dialogue systems #electronic engineering #information engineering
- Addressing Index Collapse of Large-Codebook Speech Tokenizer with Dual-Decoding Product-Quantized Variational Auto-Encoder
2024/06/05 by Haohan Guo, Guo, Haohan, Fenglong Xie +9 · 2 citations
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- Towards Data Distillation for End-to-end Spoken Conversational Question Answering
2020/10/18 by Chenyu You, Nuo Chen, You, Chenyu +7 · 1 citation
Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Signal Processing (eess.SP) #Speech and dialogue systems #Topic Modeling #electronic engineering #information engineering
- A Comparative Study of Discrete Speech Tokens for Semantic-Related Tasks with Large Language Models
2024/11/13 by Dingdong Wang, Mingyu Cui, Wang, Dingdong +7 · 2 citations
Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Sound (cs.SD) #Speech and dialogue systems #Topic Modeling #electronic engineering #information engineering
- NoreSpeech: Knowledge Distillation based Conditional Diffusion Model for Noise-robust Expressive TTS
2022/11/04 by Dongchao Yang, Songxiang Liu, Yang, Dongchao +9 · 1 citation
Computer Science · #Speech Recognition and Synthesis #Music and Audio Processing #Speech and Audio Processing
- ATRI: Mitigating Multilingual Audio Text Retrieval Inconsistencies by Reducing Data Distribution Errors
2025/02/20 by Yuguo Yin, Yuxin Xie, Yin, Yuguo +13 · 3 citations
Arts and Humanities · Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Diverse Musicological Studies #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #electronic engineering #information engineering