vix.ing · top · new · best · stats · spec

Zhikang Niu

  1. F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching
    2024/10/09 by Yushen Chen, Zhikang Niu, Chen, Yushen +14 · 2 voices · 126 citations
    Computer Science · #Music and Audio Processing #cs.SD #eess.AS
  2. MMAR: A Challenging Benchmark for Deep Reasoning in Speech, Audio, Music, and Their Mix
    2025/05/19 by Ziyang Ma, Yinghao Ma, Ma, Ziyang +62 · 49 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Multimedia (cs.MM) #Multimodal Machine Learning Applications #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #electronic engineering #information engineering
  3. SLAM-Omni: Timbre-Controllable Voice Interaction System with Single-Stage Training
    2024/12/20 by Wenxi Chen, Ziyang Ma, Chen, Wenxi +27 · 16 citations
    Computer Science · #Speech and dialogue systems #Speech Recognition and Synthesis
  4. URO-Bench: Towards Comprehensive Evaluation for End-to-End Spoken Dialogue Models
    2025/02/25 by Xiquan Li, Yan, Ruiqi, Li, Xiquan +12 · 9 citations
    Computer Science · #Topic Modeling #Speech and dialogue systems #Multimodal Machine Learning Applications
  5. EmoVoice: LLM-based Emotional Text-To-Speech Model with Freestyle Text Prompting
    2025/04/17 by Guanrou Yang, Yang, Guanrou, Yang Chen +27 · 12 citations
    Computer Science · Psychology · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #Emotion and Mood Recognition #FOS: Computer and information sciences #FOS: Electrical engineering #Mental Health via Writing #Sentiment Analysis and Opinion Mining #electronic engineering #information engineering
  6. Fast-HuBERT: An Efficient Training Framework for Self-Supervised Speech Representation Learning
    2023/09/25 by Guanrou Yang, Yang, Guanrou, Ziyang Ma +9 · 2 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and dialogue systems #Topic Modeling #electronic engineering #information engineering
  7. Accelerating Flow-Matching-Based Text-to-Speech via Empirically Pruned Step Sampling
    2025/05/26 by Yushen Chen, Zheng, Qixi, Zhikang Niu +10 · 2 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and dialogue systems #electronic engineering #information engineering
  8. NDVQ: Robust Neural Audio Codec with Normal Distribution-Based Vector Quantization
    2024/09/19 by Zhikang Niu, Niu, Zhikang, Sanyuan Chen +9 · 1 citation
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Neural Networks and Applications #Sound (cs.SD) #Speech and Audio Processing #electronic engineering #information engineering