vix.ing · top · new · best · stats · spec

Chen, Sanyuan

  1. Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers
    2023/01/05 by Wang, Chengyi, Chen, Sanyuan, Wu, Yu +10 · 133 citations
    #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
  2. BEATs: Audio Pre-Training with Acoustic Tokenizers
    2022/12/18 by Sanyuan Chen, Yu Wu, Chen, Sanyuan +11 · 84 citations
    Computer Science · #Music and Audio Processing #Speech and Audio Processing #Speech Recognition and Synthesis
  3. Movie Gen: A Cast of Media Foundation Models
    2024/10/17 by Adam Polyak, Amit Zohar, Polyak, Adam +164 · 123 citations
    Economics, Econometrics and Finance · #Cinema and Media Studies
  4. Large-scale Self-Supervised Speech Representation Learning for Automatic Speaker Verification
    2021/10/12 by Zhengyang Chen, Chen, Zhengyang, Sanyuan Chen +13 · 23 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  5. VALL-E 2: Neural Codec Language Models are Human Parity Zero-Shot Text to Speech Synthesizers
    2024/06/08 by Chen, Sanyuan, Liu, Shujie, Zhou, Long +6 · 39 citations
    #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
  6. Speak Foreign Languages with Your Own Voice: Cross-Lingual Neural Codec Language Modeling
    2023/03/07 by Ziqiang Zhang, Long Zhou, Zhang, Ziqiang +23 · 21 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Topic Modeling #electronic engineering #information engineering
  7. Meta Audiobox Aesthetics: Unified Automatic Quality Assessment for Speech, Music, and Sound
    2025/02/07 by Andros Tjandra, Tjandra, Andros, Yi-Chiao Wu +23 · 49 citations
    Computer Science · Psychology · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Multisensory perception and integration #Music Technology and Sound Studies #Music and Audio Processing #Sound (cs.SD) #electronic engineering #information engineering
  8. Recall and Learn: Fine-tuning Deep Pretrained Language Models with Less Forgetting
    2020/04/27 by Sanyuan Chen, Yutai Hou, Chen, Sanyuan +9 · 12 citations
    Computer Science · #Computation and Language (cs.CL) #Domain Adaptation and Few-Shot Learning #FOS: Computer and information sciences #Multimodal Machine Learning Applications #Topic Modeling
  9. WavLLM: Towards Robust and Adaptive Speech Large Language Model
    2024/03/31 by Shujie Hu, Hu, Shujie, Long Zhou +20 · 22 citations
    Computer Science · #Natural Language Processing Techniques #Speech Recognition and Synthesis #Topic Modeling
  10. Autoregressive Speech Synthesis without Vector Quantization
    2024/07/11 by Lingwei Meng, Meng, Lingwei, Long Zhou +21 · 23 citations
    Computer Science · #Speech Recognition and Synthesis #Speech and dialogue systems
  11. SpeechX: Neural Codec Language Model as a Versatile Speech Transformer
    2023/08/14 by Xiaofei Wang, Wang, Xiaofei, Manthan Thakker +17 · 14 citations
    Computer Science · #Speech Recognition and Synthesis #Speech and Audio Processing #Topic Modeling
  12. UniSpeech-SAT: Universal Speech Representation Learning with Speaker Aware Pre-Training
    2021/10/12 by Sanyuan Chen, Yu Wu, Chen, Sanyuan +18 · 8 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  13. VALL-E R: Robust and Efficient Zero-Shot Text-to-Speech Synthesis via Monotonic Alignment
    2024/06/12 by Han, Bing, Zhou, Long, Liu, Shujie +7 · 11 citations
    #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
  14. Why does Self-Supervised Learning for Speech Recognition Benefit Speaker Recognition?
    2022/04/27 by Chen, Sanyuan, Wu, Yu, Wang, Chengyi +8 · 4 citations
    #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
  15. Continuous Speech Separation with Conformer
    2020/08/13 by Chen, Sanyuan, Wu, Yu, Chen, Zhuo +6 · 3 citations
    #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #electronic engineering #information engineering
  16. MotherNets: Rapid Deep Ensemble Learning
    2018/09/12 by Wasay, Abdul, Hentschel, Brian, Liao, Yuze +2 · 2 citations
    #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML)
  17. SpeechLM: Enhanced Speech Pre-Training with Unpaired Textual Data
    2022/09/30 by Zhang, Ziqiang, Chen, Sanyuan, Zhou, Long +8 · 3 citations
    #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #electronic engineering #information engineering
  18. Don't shoot butterfly with rifles: Multi-channel Continuous Speech Separation with Early Exit Transformer
    2020/10/23 by Chen, Sanyuan, Wu, Yu, Chen, Zhuo +3 · 2 citations
    #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
  19. Self-Supervised Learning for speech recognition with Intermediate layer supervision
    2021/12/16 by Wang, Chengyi, Wu, Yu, Chen, Sanyuan +4 · 2 citations
    #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #electronic engineering #information engineering
  20. Investigation of Practical Aspects of Single Channel Speech Separation for ASR
    2021/07/05 by Wu, Jian, Chen, Zhuo, Chen, Sanyuan +5 · 2 citations
    #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
  21. Supervision-Guided Codebooks for Masked Prediction in Speech Pre-training
    2022/06/21 by Chengyi Wang, Wang, Chengyi, Yiming Wang +11 · 1 citation
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  22. Exploring WavLM on Speech Enhancement
    2022/11/18 by Hyungchan Song, Song, Hyungchan, Sanyuan Chen +13 · 1 citation
    Computer Science · Engineering · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Indoor and Outdoor Localization Technologies #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  23. SAM Audio: Segment Anything in Audio
    2025/12/19 by Shi, Bowen, Tjandra, Andros, Hoffman, John +11 · 3 citations
    #Audio and Speech Processing (eess.AS) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #FOS: Electrical engineering #electronic engineering #information engineering
  24. NDVQ: Robust Neural Audio Codec with Normal Distribution-Based Vector Quantization
    2024/09/19 by Zhikang Niu, Sanyuan Chen, Niu, Zhikang +9 · 1 citation
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Neural Networks and Applications #Sound (cs.SD) #Speech and Audio Processing #electronic engineering #information engineering