vix.ing · top · new · best · stats · spec

Xiquan Li

  1. MMAR: A Challenging Benchmark for Deep Reasoning in Speech, Audio, Music, and Their Mix
    2025/05/19 by Ziyang Ma, Ma, Ziyang, Yinghao Ma +62 · 49 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Multimedia (cs.MM) #Multimodal Machine Learning Applications #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #electronic engineering #information engineering
  2. SLAM-Omni: Timbre-Controllable Voice Interaction System with Single-Stage Training
    2024/12/20 by Wenxi Chen, Ziyang Ma, Chen, Wenxi +27 · 19 citations
    Computer Science · #Speech and dialogue systems #Speech Recognition and Synthesis
  3. URO-Bench: Towards Comprehensive Evaluation for End-to-End Spoken Dialogue Models
    2025/02/25 by Yan, Ruiqi, Xiquan Li, Li, Xiquan +12 · 9 citations
    Computer Science · #Topic Modeling #Speech and dialogue systems #Multimodal Machine Learning Applications
  4. SLAM-AAC: Enhancing Audio Captioning with Paraphrasing Augmentation and CLAP-Refine through LLMs
    2024/10/12 by Wenxi Chen, Chen, Wenxi, Ziyang Ma +13 · 2 citations
    Arts and Humanities · Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Multimodal Machine Learning Applications #Music and Audio Processing #Sound (cs.SD) #Subtitles and Audiovisual Media #electronic engineering #information engineering
  5. Towards Reliable Large Audio Language Model
    2025/05/25 by Ziyang Ma, Ma, Ziyang, Xiquan Li +17 · 2 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Human-Computer Interaction (cs.HC) #Multimedia (cs.MM) #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  6. Summary of DCASE 2026 Task 5: Audio-Dependent Question Answering
    2026/07/21 by Haolin He, Renhe Sun, Zheqi Dai +16
    #eess.AS