Yanmin Qian
- WavLM: Large-Scale Self-Supervised Pre-Training for Full Stack Speech Processing
2022/07/04 by Sanyuan Chen, Chengyi Wang, Zhengyang Chen +15 · 284 citations
Computer Science · #Speech Recognition and Synthesis #Speech and Audio Processing #Music and Audio Processing
- Wespeaker: A Research and Production oriented Speaker Embedding Learning Toolkit
2022/10/31 by Hongji Wang, Wang, Hongji, Chengdong Liang +13 · 38 citations
Computer Science · #Speech Recognition and Synthesis #Music and Audio Processing #Natural Language Processing Techniques
- Large-scale Self-Supervised Speech Representation Learning for Automatic Speaker Verification
2021/10/12 by Zhengyang Chen, Chen, Zhengyang, Sanyuan Chen +13 · 26 citations
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- Deep Extractor Network for Target Speaker Recovery From Single Channel Speech Mixtures
2018/07/24 by Jun Wang, Jie Chen, Wang, Jun +11 · 5 citations
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- Overview of Speaker Modeling and Its Applications: From the Lens of Deep Speaker Representation Learning
2024/07/21 by Shuai Wang, Wang, Shuai, Zhengyang Chen +7 · 9 citations
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #Speech Recognition and Synthesis #electronic engineering #information engineering
- AnoPatch: Towards Better Consistency in Machine Anomalous Sound Detection
2024/06/17 by Anbai Jiang, Bing Han, Jiang, Anbai +15 · 10 citations
Computer Science · #Music and Audio Processing #Anomaly Detection Techniques and Applications #Time Series Analysis and Forecasting
- Summary On The ICASSP 2022 Multi-Channel Multi-Party Meeting Transcription Grand Challenge
2022/02/08 by Fan Yu, Yu, Fan, Shiliang Zhang +29 · 4 citations
Computer Science · #Speech Recognition and Synthesis #Music and Audio Processing #Speech and Audio Processing
- SkiM: Skipping Memory LSTM for Low-Latency Real-Time Continuous Speech Separation
2022/01/26 by Chenda Li, Li, Chenda, Weiqin Wang +4 · 4 citations
Computer Science · Engineering · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Indoor and Outdoor Localization Technologies #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- InstructME: An Instruction Guided Music Edit And Remix Framework with Latent Diffusion Models
2023/08/28 by Bing Han, Han, Bing, Junyu Dai +15 · 5 citations
Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music Technology and Sound Studies #Music and Audio Processing #Sound (cs.SD) #Speech and Audio Processing #electronic engineering #information engineering
- Attention-based Encoder-Decoder End-to-End Neural Diarization with Embedding Enhancer
2023/09/13 by Zhengyang Chen, Chen, Zhengyang, Bing Han +5 · 5 citations
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- End-to-End Multi-speaker Speech Recognition with Transformer
2020/02/10 by Xuankai Chang, Wangyou Zhang, Chang, Xuankai +7 · 3 citations
Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- Toward Universal Speech Enhancement for Diverse Input Conditions
2023/09/29 by Wangyou Zhang, Zhang, Wangyou, Kohei Saijo +7 · 6 citations
Computer Science · Health Professions · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Infant Health and Development #Signal Processing (eess.SP) #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- Code-Switching Text Generation and Injection in Mandarin-English ASR
2023/03/20 by Haibin Yu, Yuxuan Hu, Yu, Haibin +17 · 4 citations
Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #Speech and dialogue systems #electronic engineering #information engineering
- WeSep: A Scalable and Flexible Toolkit Towards Generalizable Target Speaker Extraction
2024/09/24 by Shuai Wang, Wang, Shuai, Ke Zhang +15 · 8 citations
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- Optimizing Alignment of Speech and Language Latent Spaces for End-to-End Speech Recognition and Understanding
2021/10/23 by Wei Wang, Shuo Ren, Wang, Wei +11 · 3 citations
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #electronic engineering #information engineering
- MIMO-SPEECH: End-to-End Multi-Channel Multi-Speaker Speech Recognition
2019/10/15 by Xuankai Chang, Wangyou Zhang, Chang, Xuankai +7 · 2 citations
Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- TransVIP: Speech to Speech Translation System with Voice and Isochrony Preservation
2024/05/28 by Chenyang Le, Le, Chenyang, Yao Qian +19 · 4 citations
Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #Speech Recognition and Synthesis #electronic engineering #information engineering
- USED: Universal Speaker Extraction and Diarization
2023/09/19 by Junyi Ao, Mehmet Sinan Yıldırım, Ao, Junyi +11 · 3 citations
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- The second multi-channel multi-party meeting transcription challenge (M2MeT) 2.0): A benchmark for speaker-attributed ASR
2023/09/24 by Yuhao Liang, Liang, Yuhao, Mohan Shi +24 · 3 citations
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Topic Modeling #electronic engineering #information engineering
- Attention-based Encoder-Decoder Network for End-to-End Neural Speaker Diarization with Target Speaker Attractor
2023/05/18 by Zhengyang Chen, Chen, Zhengyang, Bing Han +5 · 2 citations
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- Self-Supervised Learning with Cluster-Aware-DINO for High-Performance Robust Speaker Verification
2023/04/12 by Bing Han, Han, Bing, Zhengyang Chen +3 · 2 citations
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- Adapting Multi-Lingual ASR Models for Handling Multiple Talkers
2023/05/30 by Chenda Li, Yao Qian, Li, Chenda +13 · 2 citations
Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Speech Recognition and Synthesis #Speech and dialogue systems #electronic engineering #information engineering
- Single-Channel Multi-talker Speech Recognition with Permutation Invariant Training
2017/07/19 by Yanmin Qian, Qian, Yanmin, Xuankai Chang +3 · 1 citation
Computer Science · Psychology · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #I.2.7 #Machine Learning (cs.LG) #Phonetics and Phonology Research #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- DQ-Whisper: Joint Distillation and Quantization for Efficient Multilingual Speech Recognition
2023/05/18 by Hang Shao, Shao, Hang, Bei Liu +7 · 2 citations
Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #Topic Modeling #electronic engineering #information engineering
- Improving Design of Input Condition Invariant Speech Enhancement
2024/01/25 by Wangyou Zhang, Jee-weon Jung, Zhang, Wangyou +5 · 2 citations
Computer Science · Engineering · #Speech and Audio Processing #Advanced Adaptive Filtering Techniques #Speech Recognition and Synthesis
- Dual-Path Modeling for Long Recording Speech Separation in Meetings
2021/02/23 by Chenda Li, Zhuo Chen, Li, Chenda +15 · 1 citation
Computer Science · #Speech and Audio Processing #Speech Recognition and Synthesis #Music and Audio Processing
- SLIDE: Integrating Speech Language Model with LLM for Spontaneous Spoken Dialogue Generation
2025/01/01 by Haitian Lu, Lu, Haitian, Gaofeng Cheng +9 · 3 citations
Computer Science · #Speech and dialogue systems #Natural Language Processing Techniques #Topic Modeling
- Diffusion-based Generative Modeling with Discriminative Guidance for Streamable Speech Enhancement
2024/06/19 by Chenda Li, Samuele Cornell, Li, Chenda +5 · 2 citations
Computer Science · #Speech and Audio Processing #Speech Recognition and Synthesis
- Self-Supervised Speaker Verification Using Dynamic Loss-Gate and Label Correction
2022/08/03 by Bing Han, Zhengyang Chen, Han, Bing +3 · 1 citation
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- CoVoMix2: Advancing Zero-Shot Dialogue Generation with Fully Non-Autoregressive Flow Matching
2025/06/01 by Yao Qian, Zhang, Leying, Qian, Yao +17 · 5 citations
Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Multimodal Machine Learning Applications #Sound (cs.SD) #Speech and dialogue systems #Topic Modeling #electronic engineering #information engineering
- Weakly-Supervised Speech Pre-training: A Case Study on Target Speech Recognition
2023/05/25 by Wangyou Zhang, Yanmin Qian, Zhang, Wangyou +1 · 1 citation
Computer Science · #Speech Recognition and Synthesis #Natural Language Processing Techniques #Music and Audio Processing
- Exploring Binary Classification Loss For Speaker Verification
2023/07/17 by Bing Han, Han, Bing, Zhengyang Chen +3 · 1 citation
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- Generating Speakers by Prompting Listener Impressions for Pre-trained Multi-Speaker Text-to-Speech Systems
2024/06/13 by Zhengyang Chen, Chen, Zhengyang, Xuechen Liu +7 · 1 citation
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and dialogue systems #electronic engineering #information engineering
- Flow-TSVAD: Target-Speaker Voice Activity Detection via Latent Flow Matching
2024/09/07 by Zhengyang Chen, Chen, Zhengyang, Bing Han +7 · 1 citation
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and dialogue systems #electronic engineering #information engineering
- Disentangling the Prosody and Semantic Information with Pre-trained Model for In-Context Learning based Zero-Shot Voice Conversion
2024/09/08 by Zhengyang Chen, Chen, Zhengyang, Shuai Wang +9 · 1 citation
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #Speech Recognition and Synthesis #electronic engineering #information engineering
- Scale This, Not That: Investigating Key Dataset Attributes for Efficient Speech Enhancement Scaling
2024/12/19 by Leying Zhang, Wangyou Zhang, Zhang, Leying +5 · 2 citations
Computer Science · Health Professions · #Speech and Audio Processing #Infant Health and Development #Speech Recognition and Synthesis
- Advanced Zero-Shot Text-to-Speech for Background Removal and Preservation with Controllable Masked Speech Prediction
2025/02/11 by Leying Zhang, Zhang, Leying, Wangyou Zhang +5 · 2 citations
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- A Data-Centric Approach to Generalizable Speech Deepfake Detection
2025/12/20 by Wen Huang, Huang, Wen, Yuchen Mao +3 · 1 citation
Computer Science · #FOS: Computer and information sciences #FOS: Electrical engineering #Generative Adversarial Networks and Image Synthesis #Hate Speech and Cyberbullying Detection #Signal Processing (eess.SP) #Sound (cs.SD) #Speech Recognition and Synthesis #electronic engineering #information engineering
- StellarTTS: Sparse Temporal Embedding for Low-Latency and Robust Speech Synthesis
2026/07/22 by Kaicheng Luo, Xuefei Gong, Yutao Sun +6
#cs.SD
- Towards Array-Invariant Speech Enhancement via Geometry-Aware Dynamic Convolution
2026/07/21 by Zhenglong Liu, Wangyou Zhang, Chenda Li +1
#eess.AS #cs.SD