vix.ing · top · new · best · stats · spec

Yanmin Qian

  1. WavLM: Large-Scale Self-Supervised Pre-Training for Full Stack Speech Processing
    2022/07/04 by Sanyuan Chen, Chengyi Wang, Zhengyang Chen +15 · 284 citations
    Computer Science · #Speech Recognition and Synthesis #Speech and Audio Processing #Music and Audio Processing
  2. Wespeaker: A Research and Production oriented Speaker Embedding Learning Toolkit
    2022/10/31 by Hongji Wang, Wang, Hongji, Chengdong Liang +13 · 38 citations
    Computer Science · #Speech Recognition and Synthesis #Music and Audio Processing #Natural Language Processing Techniques
  3. Large-scale Self-Supervised Speech Representation Learning for Automatic Speaker Verification
    2021/10/12 by Zhengyang Chen, Chen, Zhengyang, Sanyuan Chen +13 · 26 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  4. Deep Extractor Network for Target Speaker Recovery From Single Channel Speech Mixtures
    2018/07/24 by Jun Wang, Jie Chen, Wang, Jun +11 · 5 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  5. Overview of Speaker Modeling and Its Applications: From the Lens of Deep Speaker Representation Learning
    2024/07/21 by Shuai Wang, Wang, Shuai, Zhengyang Chen +7 · 9 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #Speech Recognition and Synthesis #electronic engineering #information engineering
  6. AnoPatch: Towards Better Consistency in Machine Anomalous Sound Detection
    2024/06/17 by Anbai Jiang, Bing Han, Jiang, Anbai +15 · 10 citations
    Computer Science · #Music and Audio Processing #Anomaly Detection Techniques and Applications #Time Series Analysis and Forecasting
  7. Summary On The ICASSP 2022 Multi-Channel Multi-Party Meeting Transcription Grand Challenge
    2022/02/08 by Fan Yu, Yu, Fan, Shiliang Zhang +29 · 4 citations
    Computer Science · #Speech Recognition and Synthesis #Music and Audio Processing #Speech and Audio Processing
  8. SkiM: Skipping Memory LSTM for Low-Latency Real-Time Continuous Speech Separation
    2022/01/26 by Chenda Li, Li, Chenda, Weiqin Wang +4 · 4 citations
    Computer Science · Engineering · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Indoor and Outdoor Localization Technologies #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  9. InstructME: An Instruction Guided Music Edit And Remix Framework with Latent Diffusion Models
    2023/08/28 by Bing Han, Han, Bing, Junyu Dai +15 · 5 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music Technology and Sound Studies #Music and Audio Processing #Sound (cs.SD) #Speech and Audio Processing #electronic engineering #information engineering
  10. Attention-based Encoder-Decoder End-to-End Neural Diarization with Embedding Enhancer
    2023/09/13 by Zhengyang Chen, Chen, Zhengyang, Bing Han +5 · 5 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  11. End-to-End Multi-speaker Speech Recognition with Transformer
    2020/02/10 by Xuankai Chang, Wangyou Zhang, Chang, Xuankai +7 · 3 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  12. Toward Universal Speech Enhancement for Diverse Input Conditions
    2023/09/29 by Wangyou Zhang, Zhang, Wangyou, Kohei Saijo +7 · 6 citations
    Computer Science · Health Professions · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Infant Health and Development #Signal Processing (eess.SP) #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  13. Code-Switching Text Generation and Injection in Mandarin-English ASR
    2023/03/20 by Haibin Yu, Yuxuan Hu, Yu, Haibin +17 · 4 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #Speech and dialogue systems #electronic engineering #information engineering
  14. WeSep: A Scalable and Flexible Toolkit Towards Generalizable Target Speaker Extraction
    2024/09/24 by Shuai Wang, Wang, Shuai, Ke Zhang +15 · 8 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  15. Optimizing Alignment of Speech and Language Latent Spaces for End-to-End Speech Recognition and Understanding
    2021/10/23 by Wei Wang, Shuo Ren, Wang, Wei +11 · 3 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #electronic engineering #information engineering
  16. MIMO-SPEECH: End-to-End Multi-Channel Multi-Speaker Speech Recognition
    2019/10/15 by Xuankai Chang, Wangyou Zhang, Chang, Xuankai +7 · 2 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  17. TransVIP: Speech to Speech Translation System with Voice and Isochrony Preservation
    2024/05/28 by Chenyang Le, Le, Chenyang, Yao Qian +19 · 4 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #Speech Recognition and Synthesis #electronic engineering #information engineering
  18. USED: Universal Speaker Extraction and Diarization
    2023/09/19 by Junyi Ao, Mehmet Sinan Yıldırım, Ao, Junyi +11 · 3 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  19. The second multi-channel multi-party meeting transcription challenge (M2MeT) 2.0): A benchmark for speaker-attributed ASR
    2023/09/24 by Yuhao Liang, Liang, Yuhao, Mohan Shi +24 · 3 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Topic Modeling #electronic engineering #information engineering
  20. Attention-based Encoder-Decoder Network for End-to-End Neural Speaker Diarization with Target Speaker Attractor
    2023/05/18 by Zhengyang Chen, Chen, Zhengyang, Bing Han +5 · 2 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  21. Self-Supervised Learning with Cluster-Aware-DINO for High-Performance Robust Speaker Verification
    2023/04/12 by Bing Han, Han, Bing, Zhengyang Chen +3 · 2 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  22. Adapting Multi-Lingual ASR Models for Handling Multiple Talkers
    2023/05/30 by Chenda Li, Yao Qian, Li, Chenda +13 · 2 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Speech Recognition and Synthesis #Speech and dialogue systems #electronic engineering #information engineering
  23. Single-Channel Multi-talker Speech Recognition with Permutation Invariant Training
    2017/07/19 by Yanmin Qian, Qian, Yanmin, Xuankai Chang +3 · 1 citation
    Computer Science · Psychology · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #I.2.7 #Machine Learning (cs.LG) #Phonetics and Phonology Research #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  24. DQ-Whisper: Joint Distillation and Quantization for Efficient Multilingual Speech Recognition
    2023/05/18 by Hang Shao, Shao, Hang, Bei Liu +7 · 2 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #Topic Modeling #electronic engineering #information engineering
  25. Improving Design of Input Condition Invariant Speech Enhancement
    2024/01/25 by Wangyou Zhang, Jee-weon Jung, Zhang, Wangyou +5 · 2 citations
    Computer Science · Engineering · #Speech and Audio Processing #Advanced Adaptive Filtering Techniques #Speech Recognition and Synthesis
  26. Dual-Path Modeling for Long Recording Speech Separation in Meetings
    2021/02/23 by Chenda Li, Zhuo Chen, Li, Chenda +15 · 1 citation
    Computer Science · #Speech and Audio Processing #Speech Recognition and Synthesis #Music and Audio Processing
  27. SLIDE: Integrating Speech Language Model with LLM for Spontaneous Spoken Dialogue Generation
    2025/01/01 by Haitian Lu, Lu, Haitian, Gaofeng Cheng +9 · 3 citations
    Computer Science · #Speech and dialogue systems #Natural Language Processing Techniques #Topic Modeling
  28. Diffusion-based Generative Modeling with Discriminative Guidance for Streamable Speech Enhancement
    2024/06/19 by Chenda Li, Samuele Cornell, Li, Chenda +5 · 2 citations
    Computer Science · #Speech and Audio Processing #Speech Recognition and Synthesis
  29. Self-Supervised Speaker Verification Using Dynamic Loss-Gate and Label Correction
    2022/08/03 by Bing Han, Zhengyang Chen, Han, Bing +3 · 1 citation
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  30. CoVoMix2: Advancing Zero-Shot Dialogue Generation with Fully Non-Autoregressive Flow Matching
    2025/06/01 by Yao Qian, Zhang, Leying, Qian, Yao +17 · 5 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Multimodal Machine Learning Applications #Sound (cs.SD) #Speech and dialogue systems #Topic Modeling #electronic engineering #information engineering
  31. Weakly-Supervised Speech Pre-training: A Case Study on Target Speech Recognition
    2023/05/25 by Wangyou Zhang, Yanmin Qian, Zhang, Wangyou +1 · 1 citation
    Computer Science · #Speech Recognition and Synthesis #Natural Language Processing Techniques #Music and Audio Processing
  32. Exploring Binary Classification Loss For Speaker Verification
    2023/07/17 by Bing Han, Han, Bing, Zhengyang Chen +3 · 1 citation
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  33. Generating Speakers by Prompting Listener Impressions for Pre-trained Multi-Speaker Text-to-Speech Systems
    2024/06/13 by Zhengyang Chen, Chen, Zhengyang, Xuechen Liu +7 · 1 citation
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and dialogue systems #electronic engineering #information engineering
  34. Flow-TSVAD: Target-Speaker Voice Activity Detection via Latent Flow Matching
    2024/09/07 by Zhengyang Chen, Chen, Zhengyang, Bing Han +7 · 1 citation
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and dialogue systems #electronic engineering #information engineering
  35. Disentangling the Prosody and Semantic Information with Pre-trained Model for In-Context Learning based Zero-Shot Voice Conversion
    2024/09/08 by Zhengyang Chen, Chen, Zhengyang, Shuai Wang +9 · 1 citation
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #Speech Recognition and Synthesis #electronic engineering #information engineering
  36. Scale This, Not That: Investigating Key Dataset Attributes for Efficient Speech Enhancement Scaling
    2024/12/19 by Leying Zhang, Wangyou Zhang, Zhang, Leying +5 · 2 citations
    Computer Science · Health Professions · #Speech and Audio Processing #Infant Health and Development #Speech Recognition and Synthesis
  37. Advanced Zero-Shot Text-to-Speech for Background Removal and Preservation with Controllable Masked Speech Prediction
    2025/02/11 by Leying Zhang, Zhang, Leying, Wangyou Zhang +5 · 2 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  38. A Data-Centric Approach to Generalizable Speech Deepfake Detection
    2025/12/20 by Wen Huang, Huang, Wen, Yuchen Mao +3 · 1 citation
    Computer Science · #FOS: Computer and information sciences #FOS: Electrical engineering #Generative Adversarial Networks and Image Synthesis #Hate Speech and Cyberbullying Detection #Signal Processing (eess.SP) #Sound (cs.SD) #Speech Recognition and Synthesis #electronic engineering #information engineering
  39. StellarTTS: Sparse Temporal Embedding for Low-Latency and Robust Speech Synthesis
    2026/07/22 by Kaicheng Luo, Xuefei Gong, Yutao Sun +6
    #cs.SD
  40. Towards Array-Invariant Speech Enhancement via Geometry-Aware Dynamic Convolution
    2026/07/21 by Zhenglong Liu, Wangyou Zhang, Chenda Li +1
    #eess.AS #cs.SD