Chen, Sanyuan
- Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers
2023/01/05 by Wang, Chengyi, Chen, Sanyuan, Wu, Yu +10 · 133 citations
#Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
- BEATs: Audio Pre-Training with Acoustic Tokenizers
2022/12/18 by Sanyuan Chen, Yu Wu, Chen, Sanyuan +11 · 84 citations
Computer Science · #Music and Audio Processing #Speech and Audio Processing #Speech Recognition and Synthesis
- Movie Gen: A Cast of Media Foundation Models
2024/10/17 by Adam Polyak, Amit Zohar, Polyak, Adam +164 · 123 citations
Economics, Econometrics and Finance · #Cinema and Media Studies
- Large-scale Self-Supervised Speech Representation Learning for Automatic Speaker Verification
2021/10/12 by Zhengyang Chen, Chen, Zhengyang, Sanyuan Chen +13 · 23 citations
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- VALL-E 2: Neural Codec Language Models are Human Parity Zero-Shot Text to Speech Synthesizers
2024/06/08 by Chen, Sanyuan, Liu, Shujie, Zhou, Long +6 · 39 citations
#Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
- Speak Foreign Languages with Your Own Voice: Cross-Lingual Neural Codec Language Modeling
2023/03/07 by Ziqiang Zhang, Long Zhou, Zhang, Ziqiang +23 · 21 citations
Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Topic Modeling #electronic engineering #information engineering
- Meta Audiobox Aesthetics: Unified Automatic Quality Assessment for Speech, Music, and Sound
2025/02/07 by Andros Tjandra, Tjandra, Andros, Yi-Chiao Wu +23 · 49 citations
Computer Science · Psychology · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Multisensory perception and integration #Music Technology and Sound Studies #Music and Audio Processing #Sound (cs.SD) #electronic engineering #information engineering
- Recall and Learn: Fine-tuning Deep Pretrained Language Models with Less Forgetting
2020/04/27 by Sanyuan Chen, Yutai Hou, Chen, Sanyuan +9 · 12 citations
Computer Science · #Computation and Language (cs.CL) #Domain Adaptation and Few-Shot Learning #FOS: Computer and information sciences #Multimodal Machine Learning Applications #Topic Modeling
- WavLLM: Towards Robust and Adaptive Speech Large Language Model
2024/03/31 by Shujie Hu, Hu, Shujie, Long Zhou +20 · 22 citations
Computer Science · #Natural Language Processing Techniques #Speech Recognition and Synthesis #Topic Modeling
- Autoregressive Speech Synthesis without Vector Quantization
2024/07/11 by Lingwei Meng, Meng, Lingwei, Long Zhou +21 · 23 citations
Computer Science · #Speech Recognition and Synthesis #Speech and dialogue systems
- SpeechX: Neural Codec Language Model as a Versatile Speech Transformer
2023/08/14 by Xiaofei Wang, Wang, Xiaofei, Manthan Thakker +17 · 14 citations
Computer Science · #Speech Recognition and Synthesis #Speech and Audio Processing #Topic Modeling
- UniSpeech-SAT: Universal Speech Representation Learning with Speaker Aware Pre-Training
2021/10/12 by Sanyuan Chen, Yu Wu, Chen, Sanyuan +18 · 8 citations
Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- VALL-E R: Robust and Efficient Zero-Shot Text-to-Speech Synthesis via Monotonic Alignment
2024/06/12 by Han, Bing, Zhou, Long, Liu, Shujie +7 · 11 citations
#Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
- Why does Self-Supervised Learning for Speech Recognition Benefit Speaker Recognition?
2022/04/27 by Chen, Sanyuan, Wu, Yu, Wang, Chengyi +8 · 4 citations
#Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
- Continuous Speech Separation with Conformer
2020/08/13 by Chen, Sanyuan, Wu, Yu, Chen, Zhuo +6 · 3 citations
#Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #electronic engineering #information engineering
- MotherNets: Rapid Deep Ensemble Learning
2018/09/12 by Wasay, Abdul, Hentschel, Brian, Liao, Yuze +2 · 2 citations
#FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML)
- SpeechLM: Enhanced Speech Pre-Training with Unpaired Textual Data
2022/09/30 by Zhang, Ziqiang, Chen, Sanyuan, Zhou, Long +8 · 3 citations
#Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #electronic engineering #information engineering
- Don't shoot butterfly with rifles: Multi-channel Continuous Speech Separation with Early Exit Transformer
2020/10/23 by Chen, Sanyuan, Wu, Yu, Chen, Zhuo +3 · 2 citations
#Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
- Self-Supervised Learning for speech recognition with Intermediate layer supervision
2021/12/16 by Wang, Chengyi, Wu, Yu, Chen, Sanyuan +4 · 2 citations
#Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #electronic engineering #information engineering
- Investigation of Practical Aspects of Single Channel Speech Separation for ASR
2021/07/05 by Wu, Jian, Chen, Zhuo, Chen, Sanyuan +5 · 2 citations
#Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
- Supervision-Guided Codebooks for Masked Prediction in Speech Pre-training
2022/06/21 by Chengyi Wang, Wang, Chengyi, Yiming Wang +11 · 1 citation
Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- Exploring WavLM on Speech Enhancement
2022/11/18 by Hyungchan Song, Song, Hyungchan, Sanyuan Chen +13 · 1 citation
Computer Science · Engineering · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Indoor and Outdoor Localization Technologies #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- SAM Audio: Segment Anything in Audio
2025/12/19 by Shi, Bowen, Tjandra, Andros, Hoffman, John +11 · 3 citations
#Audio and Speech Processing (eess.AS) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #FOS: Electrical engineering #electronic engineering #information engineering
- NDVQ: Robust Neural Audio Codec with Normal Distribution-Based Vector Quantization
2024/09/19 by Zhikang Niu, Sanyuan Chen, Niu, Zhikang +9 · 1 citation
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Neural Networks and Applications #Sound (cs.SD) #Speech and Audio Processing #electronic engineering #information engineering