vix.ing · top · new · best · stats · spec

Jianwei Yu

  1. Diffsound: Discrete Diffusion Model for Text-to-sound Generation
    2022/07/20 by Dongchao Yang, Jianwei Yu, Yang, Dongchao +11 · 33 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music Technology and Sound Studies #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #electronic engineering #information engineering
  2. Music Source Separation with Band-split RNN
    2022/09/30 by Yi Luo, Luo, Yi, Jianwei Yu +1 · 29 citations
    Computer Science · Engineering · #Acoustic Wave Phenomena Research #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Music and Audio Processing #Signal Processing (eess.SP) #Sound (cs.SD) #Speech and Audio Processing #electronic engineering #information engineering
  3. MMAR: A Challenging Benchmark for Deep Reasoning in Speech, Audio, Music, and Their Mix
    2025/05/19 by Ziyang Ma, Ma, Ziyang, Yinghao Ma +62 · 49 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Multimedia (cs.MM) #Multimodal Machine Learning Applications #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #electronic engineering #information engineering
  4. YuE: Scaling Open Foundation Models for Long-Form Music Generation
    2025/03/11 by Ruibin Yuan, Yuan, Ruibin, Lin, Hanfeng +110 · 28 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computer Graphics and Visualization Techniques #FOS: Computer and information sciences #FOS: Electrical engineering #Multimedia (cs.MM) #Music Technology and Sound Studies #Music and Audio Processing #Sound (cs.SD) #electronic engineering #information engineering
  5. MuQ: Self-Supervised Music Representation Learning with Mel Residual Vector Quantization
    2025/01/02 by Haina Zhu, Yizhi Zhou, Zhu, Haina +14 · 23 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  6. High Fidelity Speech Enhancement with Band-split RNN
    2022/12/01 by Jianwei Yu, Yu, Jianwei, Yi Luo +7 · 8 citations
    Computer Science · Engineering · #Advanced Adaptive Filtering Techniques #Audio and Speech Processing (eess.AS) #FOS: Electrical engineering #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  7. SECap: Speech Emotion Captioning with Large Language Model
    2023/12/16 by Yaoxun Xu, Xu, Yaoxun, Hangting Chen +15 · 9 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Topic Modeling #electronic engineering #information engineering
  8. Preference Alignment Improves Language Model-Based TTS
    2024/09/19 by Jinchuan Tian, Tian, Jinchuan, Chunlei Zhang +11 · 13 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Semantic Web and Ontologies #Speech and dialogue systems
  9. WeSep: A Scalable and Flexible Toolkit Towards Generalizable Target Speaker Extraction
    2024/09/24 by Shuai Wang, Ke Zhang, Wang, Shuai +15 · 8 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  10. Gull: A Generative Multifunctional Audio Codec
    2024/04/07 by Yi Luo, Luo, Yi, Jianwei Yu +7 · 5 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Music and Audio Processing #Signal Processing (eess.SP) #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  11. AutoPrep: An Automatic Preprocessing Framework for In-the-Wild Speech Data
    2023/09/25 by Jianwei Yu, Yu, Jianwei, Hangting Chen +15 · 4 citations
    Computer Science · #Speech Recognition and Synthesis #Natural Language Processing Techniques #Topic Modeling
  12. FRA-RIR: Fast Random Approximation of the Image-source Method
    2022/08/08 by Yi Luo, Jianwei Yu, Luo, Yi +1 · 3 citations
    Computer Science · Earth and Planetary Sciences · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Signal Processing (eess.SP) #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #Underwater Acoustics Research #electronic engineering #information engineering
  13. Bayesian Transformer Language Models for Speech Recognition
    2021/02/09 by Boyang Xue, Xue, Boyang, Jianwei Yu +15 · 2 citations
    Computer Science · #Speech Recognition and Synthesis #Topic Modeling #Natural Language Processing Techniques
  14. ASR-GLUE: A New Multi-task Benchmark for ASR-Robust Natural Language Understanding
    2021/08/30 by Lingyun Feng, Jianwei Yu, Feng, Lingyun +9 · 2 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Topic Modeling #electronic engineering #information engineering
  15. Improving Target Sound Extraction with Timestamp Information
    2022/04/02 by Helin Wang, Wang, Helin, Dongchao Yang +7 · 2 citations
    Computer Science · Engineering · #Acoustic Wave Phenomena Research #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech and Audio Processing #electronic engineering #information engineering
  16. Automatic Prosody Annotation with Pre-Trained Text-Speech Model
    2022/06/16 by Ziqian Dai, Jianwei Yu, Dai, Ziqian +13 · 2 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and dialogue systems #electronic engineering #information engineering
  17. LAE: Language-Aware Encoder for Monolingual and Multilingual ASR
    2022/06/05 by Jinchuan Tian, Jianwei Yu, Tian, Jinchuan +9 · 2 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Speech Recognition and Synthesis #Speech and Audio Processing #Speech and dialogue systems
  18. Audio-visual Recognition of Overlapped speech for the LRS2 dataset
    2020/01/06 by Jianwei Yu, Yu, Jianwei, Shi-Xiong Zhang +17 · 2 citations
    Computer Science · #Speech and Audio Processing #Music and Audio Processing #Speech Recognition and Synthesis
  19. LeVo: High-Quality Song Generation with Multi-Preference Alignment
    2025/06/09 by Shun Lei, Yanying Xu, Lei, Shun +22 · 9 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Generative Adversarial Networks and Image Synthesis #Music Technology and Sound Studies #Music and Audio Processing #Sound (cs.SD) #electronic engineering #information engineering
  20. TeCANet: Temporal-Contextual Attention Network for Environment-Aware Speech Dereverberation
    2021/03/31 by Helin Wang, Wang, Helin, Bo Wu +17 · 2 citations
    Computer Science · Neuroscience · #Speech and Audio Processing #Hearing Loss and Rehabilitation #Music and Audio Processing
  21. Deconvolutional Networks on Graph Data
    2021/10/29 by Jia Li, Jiajin Li, Li, Jia +9 · 1 citation
    Biochemistry, Genetics and Molecular Biology · Computer Science · Physics and Astronomy · #Advanced Graph Neural Networks #Bioinformatics and Genomic Networks #Complex Network Analysis Techniques #FOS: Computer and information sciences #Machine Learning (cs.LG)
  22. Mixed Precision DNN Qunatization for Overlapped Speech Separation and Recognition
    2021/11/29 by Junhao Xu, Jianwei Yu, Xu, Junhao +5 · 1 citation
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  23. NoreSpeech: Knowledge Distillation based Conditional Diffusion Model for Noise-robust Expressive TTS
    2022/11/04 by Dongchao Yang, Songxiang Liu, Yang, Dongchao +9 · 1 citation
    Computer Science · #Speech Recognition and Synthesis #Music and Audio Processing #Speech and Audio Processing
  24. Use of Speech Impairment Severity for Dysarthric Speech Recognition
    2023/05/18 by Mengzhe Geng, Zengrui Jin, Geng, Mengzhe +17 · 1 citation
    Computer Science · Medicine · Psychology · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Phonetics and Phonology Research #Sound (cs.SD) #Speech Recognition and Synthesis #Voice and Speech Disorders #electronic engineering #information engineering
  25. MuCodec: Ultra Low-Bitrate Music Codec
    2024/09/20 by Yaoxun Xu, Xu, Yaoxun, Hangting Chen +12 · 1 citation
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music Technology and Sound Studies #Music and Audio Processing #Sound (cs.SD) #Speech and Audio Processing #electronic engineering #information engineering
  26. Pushing the Frontier of Full-Song Generation: Hierarchical Autoregressive Planning Meets Flow-Matching Rendering
    2026/07/22 by Junyu Dai, Xinyue Fan, Weiqin Li +14 · 1 citation
    Computer Science · Engineering · #cs.AI #cs.SD #eess.AS
  27. VibeVoice-ASR-BitNet Technical Report
    2026/07/23 by Songchen Xu, Ting Song, Shaohan Huang +10
    #cs.SD #cs.CL #eess.AS