vix.ing · top · new · best · stats · spec

Ni, Chongjia

  1. FunAudioLLM: Voice Understanding and Generation Foundation Models for Natural Interaction Between Humans and LLMs
    2024/07/04 by An, Keyu, Chen, Qian, Deng, Chong +30 · 53 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #Speech and dialogue systems #electronic engineering #information engineering
  2. MossFormer2: Combining Transformer and RNN-Free Recurrent Network for Enhanced Time-Domain Monaural Speech Separation
    2023/12/19 by Shengkui Zhao, Zhao, Shengkui, Yukun Ma +17 · 25 citations
    Computer Science · Psychology · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Phonetics and Phonology Research #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  3. MinMo: A Multimodal Large Language Model for Seamless Voice Interaction
    2025/01/10 by Chen, Qian, Chen, Yafeng, Chen, Yanni +33 · 47 citations
    #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Human-Computer Interaction (cs.HC) #Sound (cs.SD) #electronic engineering #information engineering
  4. CosyVoice 3: Towards In-the-wild Speech Generation via Scaling-up and Post-training
    2025/05/23 by Zhihao Du, Changfeng Gao, Du, Zhihao +41 · 56 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #Speech Recognition and Synthesis #electronic engineering #information engineering
  5. InspireMusic: Integrating Super Resolution and Large Language Model for High-Fidelity Long-Form Music Generation
    2025/02/28 by Zhang, Chong, Ma, Yukun, Chen, Qian +12 · 8 citations
    #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
  6. I2CR: Improving Noise Robustness on Keyword Spotting Using Inter-Intra Contrastive Regularization
    2022/09/14 by Dianwen Ng, Jia Qi Yip, Ng, Dianwen +15 · 2 citations
    Computer Science · #Speech Recognition and Synthesis #Speech and Audio Processing #Music and Audio Processing
  7. deHuBERT: Disentangling Noise in a Self-supervised Model for Robust Speech Recognition
    2023/02/28 by Ng, Dianwen, Zhang, Ruixi, Yip, Jia Qi +7 · 2 citations
    #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
  8. Emotional Dimension Control in Language Model-Based Text-to-Speech: Spanning a Broad Spectrum of Human Emotions
    2024/09/25 by Zhou, Kun, Zhang, You, Zhao, Shengkui +9 · 4 citations
    #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
  9. FD-Bench: A Full-Duplex Benchmarking Pipeline Designed for Full Duplex Spoken Dialogue Systems
    2025/07/25 by Y Peng, Peng, Yizhou, Yi Chao +11 · 5 citations
    Computer Science · Decision Sciences · #Speech and dialogue systems #Personal Information Management and User Behavior #Topic Modeling
  10. SPGM: Prioritizing Local Features for enhanced speech separation performance
    2023/09/22 by Jia Qi Yip, Shengkui Zhao, Yip, Jia Qi +19 · 1 citation
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  11. Are Soft Prompts Good Zero-shot Learners for Speech Recognition?
    2023/09/18 by Dianwen Ng, Chong Zhang, Ng, Dianwen +17 · 1 citation
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  12. Phonetic Enhanced Language Modeling for Text-to-Speech Synthesis
    2024/06/04 by Kun Zhou, Shengkui Zhao, Zhou, Kun +17 · 2 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and dialogue systems #electronic engineering #information engineering
  13. Independent language modeling architecture for end-to-end ASR
    2019/11/25 by Van Tung Pham, Pham, Van Tung, Haihua Xu +13 · 2 citations
    Computer Science · #Speech Recognition and Synthesis #Natural Language Processing Techniques #Music and Audio Processing
  14. Fun-ASR Technical Report
    2025/09/15 by Keyu An, Yanni Chen, An, Keyu +62 · 2 citations
    Computer Science · #Speech Recognition and Synthesis #Speech and Audio Processing #Music and Audio Processing