Ziyang Ma
- Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens
2025/03/03 by Xinsheng Wang, Wang, Xinsheng, Mingqi Jiang +49 · 2 voices · 59 citations
Computer Science · #Natural Language Processing Techniques #Speech Recognition and Synthesis #Speech and dialogue systems #cs.AI #cs.SD #eess.AS
- F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching
2024/10/09 by Yushen Chen, Chen, Yushen, Zhikang Niu +14 · 2 voices · 128 citations
Computer Science · #Music and Audio Processing #cs.SD #eess.AS
- CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens
2024/07/07 by Zhihao Du, Du, Zhihao, Qian Chen +21 · 115 citations
Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Video Analysis and Summarization #electronic engineering #information engineering
- emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation
2023/12/23 by Ziyang Ma, Zhisheng Zheng, Ma, Ziyang +11 · 53 citations
Psychology · Computer Science · #Emotion and Mood Recognition #Sentiment Analysis and Opinion Mining #Speech Recognition and Synthesis
- An Embarrassingly Simple Approach for LLM with Strong ASR Capacity
2024/02/13 by Ziyang Ma, Ma, Ziyang, Guanrou Yang +19 · 25 citations
Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Fuzzy Logic and Control Systems #Multimedia (cs.MM) #Neural Networks and Applications #Sound (cs.SD) #Speech Recognition and Synthesis #electronic engineering #information engineering
- EAT: Self-Supervised Pre-Training with Efficient Audio Transformer
2024/01/07 by Wenxi Chen, Yuzhe Liang, Chen, Wenxi +7 · 22 citations
Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- MMAR: A Challenging Benchmark for Deep Reasoning in Speech, Audio, Music, and Their Mix
2025/05/19 by Ziyang Ma, Yinghao Ma, Ma, Ziyang +62 · 49 citations
Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Multimedia (cs.MM) #Multimodal Machine Learning Applications #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #electronic engineering #information engineering
- ChatMusician: Understanding and Generating Music Intrinsically with LLM
2024/02/25 by Ruibin Yuan, Yuan, Ruibin, Lin, Hanfeng +66 · 16 citations
Computer Science · Decision Sciences · Materials Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Machine Learning in Materials Science #Multimedia (cs.MM) #Scientific Computing and Data Management #Sound (cs.SD) #Topic Modeling #electronic engineering #information engineering
- Qwen3-Omni Technical Report
2025/09/22 by Xu Jin, Xu, Jin, Zhifang Guo +68 · 65 citations
Computer Science · #Speech and Audio Processing #Generative Adversarial Networks and Image Synthesis #Multimodal Machine Learning Applications
- Language Model Can Listen While Speaking
2024/08/05 by Ziyang Ma, Ma, Ziyang, Chenpeng Du +12 · 14 citations
Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Human-Computer Interaction (cs.HC) #Natural Language Processing Techniques #Sound (cs.SD) #Speech and dialogue systems #electronic engineering #information engineering
- YuE: Scaling Open Foundation Models for Long-Form Music Generation
2025/03/11 by Ruibin Yuan, Yuan, Ruibin, Lin, Hanfeng +110 · 23 citations
Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computer Graphics and Visualization Techniques #FOS: Computer and information sciences #FOS: Electrical engineering #Multimedia (cs.MM) #Music Technology and Sound Studies #Music and Audio Processing #Sound (cs.SD) #electronic engineering #information engineering
- SLAM-Omni: Timbre-Controllable Voice Interaction System with Single-Stage Training
2024/12/20 by Wenxi Chen, Chen, Wenxi, Ziyang Ma +27 · 19 citations
Computer Science · #Speech and dialogue systems #Speech Recognition and Synthesis
- MuQ: Self-Supervised Music Representation Learning with Mel Residual Vector Quantization
2025/01/02 by Haina Zhu, Yizhi Zhou, Zhu, Haina +14 · 18 citations
Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- MAP-Neo: Highly Capable and Transparent Bilingual Large Language Model Series
2024/05/29 by Ge Zhang, Zhang, Ge, Scott Qu +87 · 10 citations
Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Natural Language Processing Techniques #Speech Recognition and Synthesis #Topic Modeling
- Towards Universal Speech Discrete Tokens: A Case Study for ASR and TTS
2023/09/14 by Yifan Yang, Yang, Yifan, Feiyu Shen +11 · 8 citations
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and dialogue systems #Topic Modeling #electronic engineering #information engineering
- GigaSpeech 2: An Evolving, Large-Scale and Multi-domain ASR Corpus for Low-Resource Languages with Automated Crawling, Transcription and Refinement
2024/06/17 by Yifan Yang, Zheshu Song, Yang, Yifan +29 · 10 citations
Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Sound (cs.SD) #electronic engineering #information engineering
- MuPT: A Generative Symbolic Music Pretrained Transformer
2024/04/09 by Xingwei Qu, Qu, Xingwei, Yuelin Bai +53 · 9 citations
Computer Science · Neuroscience · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music Technology and Sound Studies #Music and Audio Processing #Neuroscience and Music Perception #Sound (cs.SD) #electronic engineering #information engineering
- Foundation Models for Music: A Survey
2024/08/26 by Yinghao Ma, Ma, Yinghao, Anders Øland +81 · 8 citations
Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Music Technology and Sound Studies #Sound (cs.SD) #electronic engineering #information engineering
- MER 2024: Semi-Supervised Learning, Noise Robustness, and Open-Vocabulary Multimodal Emotion Recognition
2024/04/26 by Zheng Lian, Lian, Zheng, Haiyang Sun +33 · 7 citations
Psychology · #Emotion and Mood Recognition #FOS: Computer and information sciences #Human-Computer Interaction (cs.HC) #Machine Learning (cs.LG)
- URO-Bench: Towards Comprehensive Evaluation for End-to-End Spoken Dialogue Models
2025/02/25 by Xiquan Li, Yan, Ruiqi, Wenxi Chen +12 · 9 citations
Computer Science · #Topic Modeling #Speech and dialogue systems #Multimodal Machine Learning Applications
- EmoVoice: LLM-based Emotional Text-To-Speech Model with Freestyle Text Prompting
2025/04/17 by Guanrou Yang, Yang Chen, Yang, Guanrou +27 · 12 citations
Computer Science · Psychology · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #Emotion and Mood Recognition #FOS: Computer and information sciences #FOS: Electrical engineering #Mental Health via Writing #Sentiment Analysis and Opinion Mining #electronic engineering #information engineering
- CTC-Assisted LLM-Based Contextual ASR
2024/11/10 by Guanrou Yang, Yang, Guanrou, Ziyang Ma +7 · 5 citations
Computer Science · #Advanced Computational Techniques and Applications #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Speech Recognition and Synthesis #electronic engineering #information engineering
- k2SSL: A Faster and Better Framework for Self-Supervised Speech Representation Learning
2024/11/26 by Yifan Yang, Yang, Yifan, Jianheng Zhuo +20 · 4 citations
Computer Science · #Speech Recognition and Synthesis #Music and Audio Processing #Speech and Audio Processing
- Chinese Tiny LLM: Pretraining a Chinese-Centric Large Language Model
2024/04/05 by Xinrun Du, Du, Xinrun, Zhouliang Yu +25 · 3 citations
Computer Science · #Natural Language Processing Techniques #Topic Modeling
- Fast-HuBERT: An Efficient Training Framework for Self-Supervised Speech Representation Learning
2023/09/25 by Guanrou Yang, Ziyang Ma, Yang, Guanrou +9 · 2 citations
Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and dialogue systems #Topic Modeling #electronic engineering #information engineering
- Pushing the Limits of Unsupervised Unit Discovery for SSL Speech Representation
2023/06/15 by Ziyang Ma, Ma, Ziyang, Zhisheng Zheng +9 · 2 citations
Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and dialogue systems #electronic engineering #information engineering
- LTCR: Long-Text Chinese Rumor Detection Dataset
2023/06/12 by Ziyang Ma, Mengsha Liu, Ma, Ziyang +5 · 2 citations
Computer Science · Social Sciences · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Misinformation and Its Impacts #Spam and Phishing Detection #Topic Modeling
- VQTalker: Towards Multilingual Talking Avatars through Facial Motion Tokenization
2024/12/13 by Tao Liu, Ziyang Ma, Liu, Tao +11 · 4 citations
Computer Science · Engineering · #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Face recognition and analysis #Human Motion and Animation #Human Pose and Action Recognition
- Enhancing Low-Resource ASR through Versatile TTS: Bridging the Data Gap
2024/10/22 by Guanrou Yang, Yu Fan, Yang, Guanrou +11 · 3 citations
Computer Science · Engineering · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Fault Detection and Control Systems #Machine Learning and ELM #Network Packet Processing and Optimization #electronic engineering #information engineering
- SLAM-AAC: Enhancing Audio Captioning with Paraphrasing Augmentation and CLAP-Refine through LLMs
2024/10/12 by Wenxi Chen, Ziyang Ma, Chen, Wenxi +13 · 2 citations
Arts and Humanities · Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Multimodal Machine Learning Applications #Music and Audio Processing #Sound (cs.SD) #Subtitles and Audiovisual Media #electronic engineering #information engineering
- Mini-Omni-Reasoner: Token-Level Thinking-in-Speaking in Large Speech Models
2025/08/18 by Zhifei Xie, Xie, Zhifei, Ziyang Ma +15 · 7 citations
Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Speech and dialogue systems #electronic engineering #information engineering
- HAM-TTS: Hierarchical Acoustic Modeling for Token-Based Zero-Shot Text-to-Speech with Model and Data Scaling
2024/03/09 by Chunhui Wang, Chang Zeng, Wang, Chunhui +15 · 1 citation
Computer Science · #Speech Recognition and Synthesis #Natural Language Processing Techniques #Topic Modeling
- Accelerating Flow-Matching-Based Text-to-Speech via Empirically Pruned Step Sampling
2025/05/26 by Yushen Chen, Zheng, Qixi, Chen, Yushen +10 · 2 citations
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and dialogue systems #electronic engineering #information engineering
- NDVQ: Robust Neural Audio Codec with Normal Distribution-Based Vector Quantization
2024/09/19 by Zhikang Niu, Niu, Zhikang, Sanyuan Chen +9 · 1 citation
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Neural Networks and Applications #Sound (cs.SD) #Speech and Audio Processing #electronic engineering #information engineering
- Towards Reliable Large Audio Language Model
2025/05/25 by Ziyang Ma, Ma, Ziyang, Xiquan Li +17 · 2 citations
Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Human-Computer Interaction (cs.HC) #Multimedia (cs.MM) #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- FVAttn: Adaptive Sparse Attention with Runtime Load Balancing for Video Generation
2026/07/17 by Hao Liu, Chenghuan Huang, Ye Huang +6
#cs.CV