vix.ing · top · new · best · stats · spec

Ziyang Ma

  1. Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens
    2025/03/03 by Xinsheng Wang, Wang, Xinsheng, Mingqi Jiang +49 · 2 voices · 59 citations
    Computer Science · #Natural Language Processing Techniques #Speech Recognition and Synthesis #Speech and dialogue systems #cs.AI #cs.SD #eess.AS
  2. F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching
    2024/10/09 by Yushen Chen, Chen, Yushen, Zhikang Niu +14 · 2 voices · 128 citations
    Computer Science · #Music and Audio Processing #cs.SD #eess.AS
  3. CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens
    2024/07/07 by Zhihao Du, Du, Zhihao, Qian Chen +21 · 115 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Video Analysis and Summarization #electronic engineering #information engineering
  4. emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation
    2023/12/23 by Ziyang Ma, Zhisheng Zheng, Ma, Ziyang +11 · 53 citations
    Psychology · Computer Science · #Emotion and Mood Recognition #Sentiment Analysis and Opinion Mining #Speech Recognition and Synthesis
  5. An Embarrassingly Simple Approach for LLM with Strong ASR Capacity
    2024/02/13 by Ziyang Ma, Ma, Ziyang, Guanrou Yang +19 · 25 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Fuzzy Logic and Control Systems #Multimedia (cs.MM) #Neural Networks and Applications #Sound (cs.SD) #Speech Recognition and Synthesis #electronic engineering #information engineering
  6. EAT: Self-Supervised Pre-Training with Efficient Audio Transformer
    2024/01/07 by Wenxi Chen, Yuzhe Liang, Chen, Wenxi +7 · 22 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  7. MMAR: A Challenging Benchmark for Deep Reasoning in Speech, Audio, Music, and Their Mix
    2025/05/19 by Ziyang Ma, Yinghao Ma, Ma, Ziyang +62 · 49 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Multimedia (cs.MM) #Multimodal Machine Learning Applications #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #electronic engineering #information engineering
  8. ChatMusician: Understanding and Generating Music Intrinsically with LLM
    2024/02/25 by Ruibin Yuan, Yuan, Ruibin, Lin, Hanfeng +66 · 16 citations
    Computer Science · Decision Sciences · Materials Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Machine Learning in Materials Science #Multimedia (cs.MM) #Scientific Computing and Data Management #Sound (cs.SD) #Topic Modeling #electronic engineering #information engineering
  9. Qwen3-Omni Technical Report
    2025/09/22 by Xu Jin, Xu, Jin, Zhifang Guo +68 · 65 citations
    Computer Science · #Speech and Audio Processing #Generative Adversarial Networks and Image Synthesis #Multimodal Machine Learning Applications
  10. Language Model Can Listen While Speaking
    2024/08/05 by Ziyang Ma, Ma, Ziyang, Chenpeng Du +12 · 14 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Human-Computer Interaction (cs.HC) #Natural Language Processing Techniques #Sound (cs.SD) #Speech and dialogue systems #electronic engineering #information engineering
  11. YuE: Scaling Open Foundation Models for Long-Form Music Generation
    2025/03/11 by Ruibin Yuan, Yuan, Ruibin, Lin, Hanfeng +110 · 23 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computer Graphics and Visualization Techniques #FOS: Computer and information sciences #FOS: Electrical engineering #Multimedia (cs.MM) #Music Technology and Sound Studies #Music and Audio Processing #Sound (cs.SD) #electronic engineering #information engineering
  12. SLAM-Omni: Timbre-Controllable Voice Interaction System with Single-Stage Training
    2024/12/20 by Wenxi Chen, Chen, Wenxi, Ziyang Ma +27 · 19 citations
    Computer Science · #Speech and dialogue systems #Speech Recognition and Synthesis
  13. MuQ: Self-Supervised Music Representation Learning with Mel Residual Vector Quantization
    2025/01/02 by Haina Zhu, Yizhi Zhou, Zhu, Haina +14 · 18 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  14. MAP-Neo: Highly Capable and Transparent Bilingual Large Language Model Series
    2024/05/29 by Ge Zhang, Zhang, Ge, Scott Qu +87 · 10 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Natural Language Processing Techniques #Speech Recognition and Synthesis #Topic Modeling
  15. Towards Universal Speech Discrete Tokens: A Case Study for ASR and TTS
    2023/09/14 by Yifan Yang, Yang, Yifan, Feiyu Shen +11 · 8 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and dialogue systems #Topic Modeling #electronic engineering #information engineering
  16. GigaSpeech 2: An Evolving, Large-Scale and Multi-domain ASR Corpus for Low-Resource Languages with Automated Crawling, Transcription and Refinement
    2024/06/17 by Yifan Yang, Zheshu Song, Yang, Yifan +29 · 10 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Sound (cs.SD) #electronic engineering #information engineering
  17. MuPT: A Generative Symbolic Music Pretrained Transformer
    2024/04/09 by Xingwei Qu, Qu, Xingwei, Yuelin Bai +53 · 9 citations
    Computer Science · Neuroscience · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music Technology and Sound Studies #Music and Audio Processing #Neuroscience and Music Perception #Sound (cs.SD) #electronic engineering #information engineering
  18. Foundation Models for Music: A Survey
    2024/08/26 by Yinghao Ma, Ma, Yinghao, Anders Øland +81 · 8 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Music Technology and Sound Studies #Sound (cs.SD) #electronic engineering #information engineering
  19. MER 2024: Semi-Supervised Learning, Noise Robustness, and Open-Vocabulary Multimodal Emotion Recognition
    2024/04/26 by Zheng Lian, Lian, Zheng, Haiyang Sun +33 · 7 citations
    Psychology · #Emotion and Mood Recognition #FOS: Computer and information sciences #Human-Computer Interaction (cs.HC) #Machine Learning (cs.LG)
  20. URO-Bench: Towards Comprehensive Evaluation for End-to-End Spoken Dialogue Models
    2025/02/25 by Xiquan Li, Yan, Ruiqi, Wenxi Chen +12 · 9 citations
    Computer Science · #Topic Modeling #Speech and dialogue systems #Multimodal Machine Learning Applications
  21. EmoVoice: LLM-based Emotional Text-To-Speech Model with Freestyle Text Prompting
    2025/04/17 by Guanrou Yang, Yang Chen, Yang, Guanrou +27 · 12 citations
    Computer Science · Psychology · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #Emotion and Mood Recognition #FOS: Computer and information sciences #FOS: Electrical engineering #Mental Health via Writing #Sentiment Analysis and Opinion Mining #electronic engineering #information engineering
  22. CTC-Assisted LLM-Based Contextual ASR
    2024/11/10 by Guanrou Yang, Yang, Guanrou, Ziyang Ma +7 · 5 citations
    Computer Science · #Advanced Computational Techniques and Applications #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Speech Recognition and Synthesis #electronic engineering #information engineering
  23. k2SSL: A Faster and Better Framework for Self-Supervised Speech Representation Learning
    2024/11/26 by Yifan Yang, Yang, Yifan, Jianheng Zhuo +20 · 4 citations
    Computer Science · #Speech Recognition and Synthesis #Music and Audio Processing #Speech and Audio Processing
  24. Chinese Tiny LLM: Pretraining a Chinese-Centric Large Language Model
    2024/04/05 by Xinrun Du, Du, Xinrun, Zhouliang Yu +25 · 3 citations
    Computer Science · #Natural Language Processing Techniques #Topic Modeling
  25. Fast-HuBERT: An Efficient Training Framework for Self-Supervised Speech Representation Learning
    2023/09/25 by Guanrou Yang, Ziyang Ma, Yang, Guanrou +9 · 2 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and dialogue systems #Topic Modeling #electronic engineering #information engineering
  26. Pushing the Limits of Unsupervised Unit Discovery for SSL Speech Representation
    2023/06/15 by Ziyang Ma, Ma, Ziyang, Zhisheng Zheng +9 · 2 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and dialogue systems #electronic engineering #information engineering
  27. LTCR: Long-Text Chinese Rumor Detection Dataset
    2023/06/12 by Ziyang Ma, Mengsha Liu, Ma, Ziyang +5 · 2 citations
    Computer Science · Social Sciences · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Misinformation and Its Impacts #Spam and Phishing Detection #Topic Modeling
  28. VQTalker: Towards Multilingual Talking Avatars through Facial Motion Tokenization
    2024/12/13 by Tao Liu, Ziyang Ma, Liu, Tao +11 · 4 citations
    Computer Science · Engineering · #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Face recognition and analysis #Human Motion and Animation #Human Pose and Action Recognition
  29. Enhancing Low-Resource ASR through Versatile TTS: Bridging the Data Gap
    2024/10/22 by Guanrou Yang, Yu Fan, Yang, Guanrou +11 · 3 citations
    Computer Science · Engineering · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Fault Detection and Control Systems #Machine Learning and ELM #Network Packet Processing and Optimization #electronic engineering #information engineering
  30. SLAM-AAC: Enhancing Audio Captioning with Paraphrasing Augmentation and CLAP-Refine through LLMs
    2024/10/12 by Wenxi Chen, Ziyang Ma, Chen, Wenxi +13 · 2 citations
    Arts and Humanities · Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Multimodal Machine Learning Applications #Music and Audio Processing #Sound (cs.SD) #Subtitles and Audiovisual Media #electronic engineering #information engineering
  31. Mini-Omni-Reasoner: Token-Level Thinking-in-Speaking in Large Speech Models
    2025/08/18 by Zhifei Xie, Xie, Zhifei, Ziyang Ma +15 · 7 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Speech and dialogue systems #electronic engineering #information engineering
  32. HAM-TTS: Hierarchical Acoustic Modeling for Token-Based Zero-Shot Text-to-Speech with Model and Data Scaling
    2024/03/09 by Chunhui Wang, Chang Zeng, Wang, Chunhui +15 · 1 citation
    Computer Science · #Speech Recognition and Synthesis #Natural Language Processing Techniques #Topic Modeling
  33. Accelerating Flow-Matching-Based Text-to-Speech via Empirically Pruned Step Sampling
    2025/05/26 by Yushen Chen, Zheng, Qixi, Chen, Yushen +10 · 2 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and dialogue systems #electronic engineering #information engineering
  34. NDVQ: Robust Neural Audio Codec with Normal Distribution-Based Vector Quantization
    2024/09/19 by Zhikang Niu, Niu, Zhikang, Sanyuan Chen +9 · 1 citation
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Neural Networks and Applications #Sound (cs.SD) #Speech and Audio Processing #electronic engineering #information engineering
  35. Towards Reliable Large Audio Language Model
    2025/05/25 by Ziyang Ma, Ma, Ziyang, Xiquan Li +17 · 2 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Human-Computer Interaction (cs.HC) #Multimedia (cs.MM) #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  36. FVAttn: Adaptive Sparse Attention with Runtime Load Balancing for Video Generation
    2026/07/17 by Hao Liu, Chenghuan Huang, Ye Huang +6
    #cs.CV