vix.ing · top · new · best · stats · spec

Lingwei Meng

  1. Autoregressive Speech Synthesis without Vector Quantization
    2024/07/11 by Lingwei Meng, Long Zhou, Meng, Lingwei +21 · 23 citations
    Computer Science · #Speech Recognition and Synthesis #Speech and dialogue systems
  2. ARLON: Boosting Diffusion Transformers with Autoregressive Models for Long Video Generation
    2024/10/27 by Zongyi Li, Li, Zongyi, Shujie Hu +16 · 13 citations
    Computer Science · #Generative Adversarial Networks and Image Synthesis #Advanced Image Processing Techniques #Image and Video Quality Assessment
  3. The defender's perspective on automatic speaker verification: An overview
    2023/05/22 by Haibin Wu, Wu, Haibin, Jiawen Kang +7 · 4 citations
    Computer Science · #Adversarial Robustness in Machine Learning #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Network Security and Intrusion Detection #Sound (cs.SD) #Speech Recognition and Synthesis #electronic engineering #information engineering
  4. Empowering Whisper as a Joint Multi-Talker and Target-Talker Speech Recognition System
    2024/07/13 by Lingwei Meng, Meng, Lingwei, Jiawen Kang +11 · 5 citations
    Computer Science · Environmental Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #Educational Reforms and Innovations #FOS: Computer and information sciences #FOS: Electrical engineering #Hate Speech and Cyberbullying Detection #Sound (cs.SD) #electronic engineering #information engineering
  5. Spoofing-Aware Speaker Verification by Multi-Level Fusion
    2022/03/29 by Haibin Wu, Wu, Haibin, Lingwei Meng +13 · 2 citations
    Computer Science · Medicine · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #Voice and Speech Disorders #electronic engineering #information engineering
  6. Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey
    2024/12/16 by Liang Chen, Zekun Wang, Chen, Liang +49 · 5 citations
    Computer Science · #Natural Language Processing Techniques #Topic Modeling
  7. Unified Modeling of Multi-Talker Overlapped Speech Recognition and Diarization with a Sidecar Separator
    2023/05/25 by Lingwei Meng, Jiawen Kang, Meng, Lingwei +9 · 2 citations
    Computer Science · Psychology · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Phonetics and Phonology Research #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  8. LibriheavyMix: A 20,000-Hour Dataset for Single-Channel Reverberant Multi-Talker Speech Separation, ASR and Speaker Diarization
    2024/09/01 by Zengrui Jin, Jin, Zengrui, Yifan Yang +22 · 2 citations
    Computer Science · Health Professions · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Infant Health and Development #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  9. UNIT-DSR: Dysarthric Speech Reconstruction System Using Speech Unit Normalization
    2024/01/26 by Yuejiao Wang, Xixin Wu, Wang, Yuejiao +7 · 1 citation
    Computer Science · Medicine · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #Voice and Speech Disorders #electronic engineering #information engineering
  10. Pseudo-Autoregressive Neural Codec Language Models for Efficient Zero-Shot Text-to-Speech Synthesis
    2025/04/14 by Yifan Yang, Shujie Liu, Yang, Yifan +22 · 4 citations
    Computer Science · #Speech Recognition and Synthesis #Natural Language Processing Techniques #Topic Modeling
  11. StreamMel: Real-Time Zero-shot Text-to-Speech via Interleaved Continuous Autoregressive Modeling
    2025/06/14 by Hui Wang, Wang, Hui, Shujie Liu +16 · 4 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Topic Modeling #electronic engineering #information engineering
  12. Disentangling Speakers in Multi-Talker Speech Recognition with Speaker-Aware CTC
    2024/09/19 by Jiawen Kang, Lingwei Meng, Kang, Jiawen +11 · 1 citation
    Computer Science · #Advanced Data Compression Techniques #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  13. Dynamic Uncertainty Learning with Noisy Correspondence for Text-Based Person Search
    2025/05/10 by Zequn Xie, Xie, Zequn, H. Ji +5 · 2 citations
    Medicine · Decision Sciences · #Data-Driven Disease Surveillance #Data Quality and Management