Xie Chen
- Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens
2025/03/03 by Xinsheng Wang, Mingqi Jiang, Ming Jiang +49 · 2 voices · 72 citations
Computer Science · #Natural Language Processing Techniques #Speech Recognition and Synthesis #Speech and dialogue systems #cs.AI #cs.SD #eess.AS
- F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching
2024/10/09 by Yushen Chen, Zhikang Niu, Chen, Yushen +14 · 2 voices · 149 citations
Computer Science · #Music and Audio Processing #cs.SD #eess.AS
- emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation
2023/12/23 by Ziyang Ma, Ma, Ziyang, Zhisheng Zheng +11 · 60 citations
Psychology · Computer Science · #Emotion and Mood Recognition #Sentiment Analysis and Opinion Mining #Speech Recognition and Synthesis
- Fracton phases of matter
2020/02/29 by Michael Pretko, Xie Chen, Yizhi You · 25 citations
Materials Science · Physics and Astronomy · #Chemical and Physical Properties of Materials #Quantum many-body systems #Topological Materials and Phenomena
- EAT: Self-Supervised Pre-Training with Efficient Audio Transformer
2024/01/07 by Wenxi Chen, Yuzhe Liang, Chen, Wenxi +7 · 32 citations
Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- An Embarrassingly Simple Approach for LLM with Strong ASR Capacity
2024/02/13 by Ziyang Ma, Guanrou Yang, Ma, Ziyang +19 · 31 citations
Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Fuzzy Logic and Control Systems #Multimedia (cs.MM) #Neural Networks and Applications #Sound (cs.SD) #Speech Recognition and Synthesis #electronic engineering #information engineering
- AniTalker: Animate Vivid and Diverse Talking Faces through Identity-Decoupled Facial Motion Encoding
2024/05/06 by Tao Liu, Liu, Tao, Feilong Chen +11 · 22 citations
Computer Science · #Artificial Intelligence (cs.AI) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Face recognition and analysis
- BAT: Learning to Reason about Spatial Sounds with Large Language Models
2024/02/02 by Zhisheng Zheng, Zheng, Zhisheng, Puyuan Peng +9 · 18 citations
Computer Science · Social Sciences · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #Data Management and Algorithms #FOS: Computer and information sciences #FOS: Electrical engineering #Geographic Information Systems Studies #Sound (cs.SD) #Speech and dialogue systems #electronic engineering #information engineering
- Language Model Can Listen While Speaking
2024/08/05 by Ziyang Ma, Ma, Ziyang, Chenpeng Du +12 · 17 citations
Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Human-Computer Interaction (cs.HC) #Natural Language Processing Techniques #Sound (cs.SD) #Speech and dialogue systems #electronic engineering #information engineering
- YuE: Scaling Open Foundation Models for Long-Form Music Generation
2025/03/11 by Ruibin Yuan, Yuan, Ruibin, Shuyue Guo +110 · 28 citations
Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computer Graphics and Visualization Techniques #FOS: Computer and information sciences #FOS: Electrical engineering #Multimedia (cs.MM) #Music Technology and Sound Studies #Music and Audio Processing #Sound (cs.SD) #electronic engineering #information engineering
- MuQ: Self-Supervised Music Representation Learning with Mel Residual Vector Quantization
2025/01/02 by Haina Zhu, Zhu, Haina, Yizhi Zhou +14 · 23 citations
Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- SLAM-Omni: Timbre-Controllable Voice Interaction System with Single-Stage Training
2024/12/20 by Wenxi Chen, Chen, Wenxi, Ziyang Ma +27 · 24 citations
Computer Science · #Speech and dialogue systems #Speech Recognition and Synthesis
- GigaSpeech 2: An Evolving, Large-Scale and Multi-domain ASR Corpus for Low-Resource Languages with Automated Crawling, Transcription and Refinement
2024/06/17 by Yifan Yang, Yang, Yifan, Zheshu Song +29 · 14 citations
Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Sound (cs.SD) #electronic engineering #information engineering
- EmoDiff: Intensity Controllable Emotional Text-to-Speech with Soft-Label Guidance
2022/11/17 by Yiwei Guo, Chenpeng Du, Guo, Yiwei +5 · 8 citations
Computer Science · Psychology · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Human-Computer Interaction (cs.HC) #Machine Learning (cs.LG) #Phonetics and Phonology Research #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- Towards Universal Speech Discrete Tokens: A Case Study for ASR and TTS
2023/09/14 by Yifan Yang, Yang, Yifan, Feiyu Shen +11 · 9 citations
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and dialogue systems #Topic Modeling #electronic engineering #information engineering
- AnoPatch: Towards Better Consistency in Machine Anomalous Sound Detection
2024/06/17 by Anbai Jiang, Bing Han, Jiang, Anbai +15 · 10 citations
Computer Science · #Music and Audio Processing #Anomaly Detection Techniques and Applications #Time Series Analysis and Forecasting
- MER 2024: Semi-Supervised Learning, Noise Robustness, and Open-Vocabulary Multimodal Emotion Recognition
2024/04/26 by Zheng Lian, Haiyang Sun, Lian, Zheng +33 · 9 citations
Psychology · #Emotion and Mood Recognition #FOS: Computer and information sciences #Human-Computer Interaction (cs.HC) #Machine Learning (cs.LG)
- LSCodec: Low-Bitrate and Speaker-Decoupled Discrete Speech Codec
2024/10/21 by Yiwei Guo, Zhihan Li, Guo, Yiwei +9 · 9 citations
Computer Science · #Advanced Data Compression Techniques #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- URO-Bench: Towards Comprehensive Evaluation for End-to-End Spoken Dialogue Models
2025/02/25 by Xiquan Li, Yan, Ruiqi, Wenxi Chen +12 · 9 citations
Computer Science · #Topic Modeling #Speech and dialogue systems #Multimodal Machine Learning Applications
- EmoVoice: LLM-based Emotional Text-To-Speech Model with Freestyle Text Prompting
2025/04/17 by Guanrou Yang, Yang, Guanrou, Yang Chen +27 · 16 citations
Computer Science · Psychology · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #Emotion and Mood Recognition #FOS: Computer and information sciences #FOS: Electrical engineering #Mental Health via Writing #Sentiment Analysis and Opinion Mining #electronic engineering #information engineering
- Sequential Adiabatic Generation of Chiral Topological States
2024/02/05 by Xie Chen, Michael Hermele, Chen, Xie +3 · 4 citations
Physics and Astronomy · Computer Science · #Spectroscopy and Quantum Chemical Studies #Nonlinear Dynamics and Pattern Formation
- The Interspeech 2024 Challenge on Speech Processing Using Discrete Units
2024/06/11 by Xuankai Chang, Chang, Xuankai, Jiatong Shi +17 · 5 citations
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #Speech Recognition and Synthesis #electronic engineering #information engineering
- CTC-Assisted LLM-Based Contextual ASR
2024/11/10 by Guanrou Yang, Yang, Guanrou, Ziyang Ma +7 · 6 citations
Computer Science · #Advanced Computational Techniques and Applications #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Speech Recognition and Synthesis #electronic engineering #information engineering
- VQTalker: Towards Multilingual Talking Avatars through Facial Motion Tokenization
2024/12/13 by Tao Liu, Liu, Tao, Ziyang Ma +11 · 6 citations
Computer Science · Engineering · #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Face recognition and analysis #Human Motion and Animation #Human Pose and Action Recognition
- Factorized Neural Transducer for Efficient Language Model Adaptation
2021/09/27 by Xie Chen, Chen, Xie, Zhong Meng +5 · 2 citations
Computer Science · #Speech Recognition and Synthesis #Topic Modeling #Music and Audio Processing
- SEF-VC: Speaker Embedding Free Zero-Shot Voice Conversion with Cross Attention
2023/12/14 by Junjie Li, Yiwei Guo, Li, Junjie +5 · 3 citations
Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- DRCap: Decoding CLAP Latents with Retrieval-Augmented Generation for Zero-shot Audio Captioning
2024/10/12 by Xiquan Li, Li, Xiquan, Wenxi Chen +13 · 5 citations
Arts and Humanities · Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Natural Language Processing Techniques #Sound (cs.SD) #Subtitles and Audiovisual Media #electronic engineering #information engineering
- k2SSL: A Faster and Better Framework for Self-Supervised Speech Representation Learning
2024/11/26 by Yifan Yang, Jianheng Zhuo, Yang, Yifan +20 · 4 citations
Computer Science · #Speech Recognition and Synthesis #Music and Audio Processing #Speech and Audio Processing
- Fast-HuBERT: An Efficient Training Framework for Self-Supervised Speech Representation Learning
2023/09/25 by Guanrou Yang, Ziyang Ma, Yang, Guanrou +9 · 2 citations
Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and dialogue systems #Topic Modeling #electronic engineering #information engineering
- Pushing the Limits of Unsupervised Unit Discovery for SSL Speech Representation
2023/06/15 by Ziyang Ma, Zhisheng Zheng, Ma, Ziyang +9 · 2 citations
Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and dialogue systems #electronic engineering #information engineering
- Enhancing Low-Resource ASR through Versatile TTS: Bridging the Data Gap
2024/10/22 by Guanrou Yang, Yu Fan, Yang, Guanrou +11 · 4 citations
Computer Science · Engineering · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Fault Detection and Control Systems #Machine Learning and ELM #Network Packet Processing and Optimization #electronic engineering #information engineering
- Attention-Constrained Inference for Robust Decoder-Only Text-to-Speech
2024/04/30 by Hankun Wang, Chenpeng Du, Wang, Hankun +9 · 2 citations
Computer Science · #Speech Recognition and Synthesis
- The Symmetry Taco: Equivalences between Gapped, Gapless, and Mixed-State SPTs
2025/07/07 by Marvin Qi, Qi, Marvin, Ramanjit Sohal +7 · 6 citations
Physics and Astronomy · #FOS: Physical sciences #Force Microscopy Techniques and Applications #High Energy Physics - Theory (hep-th) #Quantum Physics (quant-ph) #Strongly Correlated Electrons (cond-mat.str-el)
- LSTM-LM with Long-Term History for First-Pass Decoding in Conversational Speech Recognition
2020/10/21 by Xie Chen, Sarangarajan Parthasarathy, Chen, Xie +7 · 1 citation
Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Speech Recognition and Synthesis #Speech and dialogue systems #Topic Modeling
- Exploring SSL Discrete Tokens for Multilingual ASR
2024/09/13 by Mingyu Cui, Daxin Tan, Cui, Mingyu +12 · 2 citations
Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and dialogue systems #electronic engineering #information engineering
- SLAM-AAC: Enhancing Audio Captioning with Paraphrasing Augmentation and CLAP-Refine through LLMs
2024/10/12 by Wenxi Chen, Ziyang Ma, Chen, Wenxi +13 · 2 citations
Arts and Humanities · Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Multimodal Machine Learning Applications #Music and Audio Processing #Sound (cs.SD) #Subtitles and Audiovisual Media #electronic engineering #information engineering
- Long-span language modeling for speech recognition
2019/11/11 by Sarangarajan Parthasarathy, Parthasarathy, Sarangarajan, William A. Gale +7 · 2 citations
Computer Science · #Speech Recognition and Synthesis #Topic Modeling #Natural Language Processing Techniques
- Exploring Effective Fusion Algorithms for Speech Based Self-Supervised Learning Models
2022/12/20 by Changli Tang, Tang, Changli, Yujin Wang +5 · 1 citation
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #electronic engineering #information engineering
- Blank-regularized CTC for Frame Skipping in Neural Transducer
2023/05/19 by Yifan Yang, Xiaoyu Yang, Yang, Yifan +14 · 1 citation
Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Neural Networks and Applications #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- Ground State Degeneracy of Infinite-Component Chern-Simons-Maxwell Theories: Foliated vs. Non-foliated Fracton Orders
2023/06/01 by Xie Chen, Chen, Xie, Ho Tat Lam +3 · 1 citation
Physics and Astronomy · Mathematics · #Theoretical and Computational Physics #Algebraic structures and combinatorial models #Quantum many-body systems
- DSE-TTS: Dual Speaker Embedding for Cross-Lingual Text-to-Speech
2023/06/25 by Sen Liu, Yiwei Guo, Liu, Sen +7 · 1 citation
Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Topic Modeling #electronic engineering #information engineering
- Pseudo-Autoregressive Neural Codec Language Models for Efficient Zero-Shot Text-to-Speech Synthesis
2025/04/14 by Yifan Yang, Yang, Yifan, Shujie Liu +22 · 6 citations
Computer Science · #Speech Recognition and Synthesis #Natural Language Processing Techniques #Topic Modeling
- Acoustic BPE for Speech Generation with Discrete Tokens
2023/10/23 by Feiyu Shen, Shen, Feiyu, Yiwei Guo +7 · 1 citation
Computer Science · #Speech Recognition and Synthesis #Speech and dialogue systems #Topic Modeling
- Accelerating Flow-Matching-Based Text-to-Speech via Empirically Pruned Step Sampling
2025/05/26 by Zheng, Qixi, Yushen Chen, Chen, Yushen +10 · 2 citations
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and dialogue systems #electronic engineering #information engineering
- vec2wav 2.0: Advancing Voice Conversion via Discrete Token Vocoders
2024/09/03 by Yiwei Guo, Zhihan Li, Guo, Yiwei +13 · 1 citation
Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- Towards Reliable Large Audio Language Model
2025/05/25 by Ziyang Ma, Ma, Ziyang, Xiquan Li +17 · 2 citations
Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Human-Computer Interaction (cs.HC) #Multimedia (cs.MM) #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- Towards Effective and Compact Contextual Representation for Conformer Transducer Speech Recognition Systems
2023/06/23 by Mingyu Cui, Cui, Mingyu, Jiawen Kang +11 · 1 citation
Computer Science · Psychology · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Phonetics and Phonology Research #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- SimulS2S-LLM: Unlocking Simultaneous Inference of Speech LLMs for Speech-to-Speech Translation
2025/04/22 by Keqi Deng, Deng, Keqi, Wenxi Chen +5 · 2 citations
Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Topic Modeling #electronic engineering #information engineering
- Fusion of Low-Entanglement Excitations in 2D Toric Code
2024/09/11 by Jingyu Zhao, Xie Chen, Zhao, Jing-Yu +1 · 1 citation
Computer Science · Physics and Astronomy · #FOS: Physical sciences #Quantum Computing Algorithms and Architecture #Quantum Information and Cryptography #Quantum and electron transport phenomena #Strongly Correlated Electrons (cond-mat.str-el)
- OmniVAE: An Audio-Video VAE with Cross-Modal Alignment for Joint Generation
2026/07/26 by Jun Zhan, Chen Yang, Yitian Gong +23
#cs.SD #cs.CV
- AgenticASR: Refining Speech Recognition in Real-World Scenarios via an Agentic Approach
2026/07/30 by Zixuan Jiang, Binghao Qiang, Jiaying Chi +3
Computer Science · #cs.AI
- Efficient Chain-of-Modality Reasoning via Progressive Compression for Spoken Language Models
2026/07/22 by Pengchao Feng, Chao-Hong Tan, Qian Chen +3
#cs.CL #cs.SD
- Summary of DCASE 2026 Task 5: Audio-Dependent Question Answering
2026/07/21 by Haolin He, Renhe Sun, Zheqi Dai +16
#eess.AS
- X-Translator: A Real-Time Multilingual Speaker-Aware Speech-to-Speech Translation System
2026/07/20 by Yuxiang Zhao, Yichi Zhang, Yanjie An +10
#eess.AS