vix.ing · top · new · best · stats · spec

Tianrui Wang

  1. On decoder-only architecture for speech-to-text and large language model integration
    2023/07/08 by Jian Wu, Yashesh Gaur, Wu, Jian +19 · 31 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Topic Modeling #electronic engineering #information engineering
  2. MMAR: A Challenging Benchmark for Deep Reasoning in Speech, Audio, Music, and Their Mix
    2025/05/19 by Ziyang Ma, Yinghao Ma, Ma, Ziyang +62 · 50 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Multimedia (cs.MM) #Multimodal Machine Learning Applications #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #electronic engineering #information engineering
  3. EmoVoice: LLM-based Emotional Text-To-Speech Model with Freestyle Text Prompting
    2025/04/17 by Guanrou Yang, Yang Chen, Yang, Guanrou +27 · 16 citations
    Computer Science · Psychology · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #Emotion and Mood Recognition #FOS: Computer and information sciences #FOS: Electrical engineering #Mental Health via Writing #Sentiment Analysis and Opinion Mining #electronic engineering #information engineering
  4. Mamba-SEUNet: Mamba UNet for Monaural Speech Enhancement
    2024/12/21 by Junyu Wang, Wang, Junyu, Tianrui Wang +8 · 6 citations
    Computer Science · Psychology · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Phonetics and Phonology Research #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  5. VQ-CTAP: Cross-Modal Fine-Grained Sequence Representation Learning for Speech Processing
    2024/08/11 by Chunyu Qiang, Qiang, Chunyu, Geng Wang +26 · 6 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and dialogue systems #electronic engineering #information engineering
  6. VocalRender: Score-Native Singing Voice Synthesis for Real-World Composition
    2026/07/30 by Yukun Chen, Tianrui Wang, Zhaoxi Mu +2
    Computer Science · #cs.SD #cs.AI