vix.ing · top · new · best · stats · spec

Dongchao Yang

  1. AudioGPT: Understanding and Generating Speech, Music, Sound, and Talking Head
    2023/04/25 by Rongjie Huang, Huang, Rongjie, Mingze Li +23 · 1 voice · 52 citations
    Computer Science · #Music and Audio Processing #Speech Recognition and Synthesis #Topic Modeling
  2. NaturalSpeech 3: Zero-Shot Speech Synthesis with Factorized Codec and Diffusion Models
    2024/03/05 by Zeqian Ju, Ju, Zeqian, Yuancheng Wang +35 · 75 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  3. Make-An-Audio: Text-To-Audio Generation with Prompt-Enhanced Diffusion Models
    2023/01/30 by Rongjie Huang, Jiawei Huang, Huang, Rongjie +17 · 42 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Generative Adversarial Networks and Image Synthesis #Machine Learning (cs.LG) #Multimedia (cs.MM) #Music Technology and Sound Studies #Music and Audio Processing #Sound (cs.SD) #electronic engineering #information engineering
  4. Diffsound: Discrete Diffusion Model for Text-to-sound Generation
    2022/07/20 by Dongchao Yang, Yang, Dongchao, Jianwei Yu +11 · 26 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music Technology and Sound Studies #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #electronic engineering #information engineering
  5. HiFi-Codec: Group-residual Vector quantization for High Fidelity Audio Codec
    2023/05/04 by Dongchao Yang, Yang, Dongchao, Songxiang Liu +9 · 31 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  6. Make-An-Audio 2: Temporal-Enhanced Text-to-Audio Generation
    2023/05/29 by Jiawei Huang, Huang, Jiawei, Yi Ren +17 · 21 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Multimedia (cs.MM) #Music Technology and Sound Studies #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #electronic engineering #information engineering
  7. InstructTTS: Modelling Expressive TTS in Discrete Latent Space with Natural Language Style Prompt
    2023/01/31 by Dongchao Yang, Yang, Dongchao, Songxiang Liu +7 · 14 citations
    Computer Science · #Speech Recognition and Synthesis #Speech and Audio Processing #Music and Audio Processing
  8. MMSU: A Massive Multi-task Spoken Language Understanding and Reasoning Benchmark
    2025/06/05 by Dingdong Wang, Wang, Dingdong, Junan Li +11 · 31 citations
    Computer Science · Psychology · #Speech Recognition and Synthesis #Emotion and Mood Recognition #Speech and dialogue systems
  9. UniAudio 1.5: Large Language Model-driven Audio Codec is A Few-shot Audio Task Learner
    2024/06/14 by Dongchao Yang, Yang, Dongchao, Haohan Guo +13 · 10 citations
    Computer Science · #Music and Audio Processing #Speech Recognition and Synthesis #Speech and Audio Processing
  10. PromptTTS 2: Describing and Generating Voices with Text Prompt
    2023/09/05 by Yichong Leng, Zhifang Guo, Leng, Yichong +27 · 7 citations
    Computer Science · #Speech Recognition and Synthesis #Topic Modeling #Natural Language Processing Techniques
  11. SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline
    2025/05/25 by Helin Wang, Jiarui Hai, Wang, Helin +17 · 1 voice · 4 citations
    Computer Science · Engineering · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #cs.AI #cs.SD #eess.AS #electronic engineering #information engineering
  12. Make-A-Voice: Unified Voice Synthesis With Discrete Representation
    2023/05/30 by Rongjie Huang, Huang, Rongjie, Chunlei Zhang +17 · 5 citations
    Computer Science · #Speech Recognition and Synthesis #Music and Audio Processing #Topic Modeling
  13. RALL-E: Robust Codec Language Modeling with Chain-of-Thought Prompting for Text-to-Speech Synthesis
    2024/04/04 by Detai Xin, Xu Tan, Xin, Detai +19 · 4 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Topic Modeling #electronic engineering #information engineering
  14. SimpleSpeech 2: Towards Simple and Efficient Text-to-Speech with Flow-based Scalar Latent Transformer Diffusion Models
    2024/08/25 by Dongchao Yang, Rongjie Huang, Yang, Dongchao +13 · 5 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Topic Modeling #electronic engineering #information engineering
  15. InSerter: Speech Instruction Following with Unsupervised Interleaved Pre-training
    2025/03/04 by Dingdong Wang, Wang, Dingdong, Jin Xu +14 · 7 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Human-Computer Interaction (cs.HC) #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Topic Modeling #electronic engineering #information engineering
  16. Improving Target Sound Extraction with Timestamp Information
    2022/04/02 by Helin Wang, Dongchao Yang, Wang, Helin +7 · 2 citations
    Computer Science · Engineering · #Acoustic Wave Phenomena Research #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech and Audio Processing #electronic engineering #information engineering
  17. AudioComposer: Towards Fine-grained Audio Generation with Natural Language Descriptions
    2024/09/19 by Yuanyuan Wang, Wang, Yuanyuan, Hangting Chen +7 · 4 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music Technology and Sound Studies #Music and Audio Processing #Sound (cs.SD) #Speech and Audio Processing #electronic engineering #information engineering
  18. ALMTokenizer: A Low-bitrate and Semantic-rich Audio Codec Tokenizer for Audio Language Modeling
    2025/04/14 by Dongchao Yang, Songxiang Liu, Yang, Dongchao +21 · 8 citations
    Computer Science · #FOS: Computer and information sciences #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing
  19. VARGPT-v1.1: Improve Visual Autoregressive Large Unified Model via Iterative Instruction Tuning and Reinforcement Learning
    2025/04/03 by Xianwei Zhuang, Yuxin Xie, Zhuang, Xianwei +13 · 7 citations
    Computer Science · Biochemistry, Genetics and Molecular Biology · #Generative Adversarial Networks and Image Synthesis #Multimodal Machine Learning Applications #Cell Image Analysis Techniques
  20. SoCodec: A Semantic-Ordered Multi-Stream Speech Codec for Efficient Language Model Based Text-to-Speech Synthesis
    2024/09/02 by Haohan Guo, Guo, Haohan, Fenglong Xie +11 · 3 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and dialogue systems #electronic engineering #information engineering
  21. Addressing Index Collapse of Large-Codebook Speech Tokenizer with Dual-Decoding Product-Quantized Variational Auto-Encoder
    2024/06/05 by Haohan Guo, Guo, Haohan, Fenglong Xie +9 · 2 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  22. Towards Data Distillation for End-to-end Spoken Conversational Question Answering
    2020/10/18 by Chenyu You, Nuo Chen, You, Chenyu +7 · 1 citation
    Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Signal Processing (eess.SP) #Speech and dialogue systems #Topic Modeling #electronic engineering #information engineering
  23. A Comparative Study of Discrete Speech Tokens for Semantic-Related Tasks with Large Language Models
    2024/11/13 by Dingdong Wang, Mingyu Cui, Wang, Dingdong +7 · 2 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Sound (cs.SD) #Speech and dialogue systems #Topic Modeling #electronic engineering #information engineering
  24. NoreSpeech: Knowledge Distillation based Conditional Diffusion Model for Noise-robust Expressive TTS
    2022/11/04 by Dongchao Yang, Songxiang Liu, Yang, Dongchao +9 · 1 citation
    Computer Science · #Speech Recognition and Synthesis #Music and Audio Processing #Speech and Audio Processing
  25. ATRI: Mitigating Multilingual Audio Text Retrieval Inconsistencies by Reducing Data Distribution Errors
    2025/02/20 by Yuguo Yin, Yuxin Xie, Yin, Yuguo +13 · 3 citations
    Arts and Humanities · Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Diverse Musicological Studies #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #electronic engineering #information engineering