Guo, Pengcheng
- WenetSpeech: A 10000+ Hours Multi-domain Mandarin Corpus for Speech Recognition
2021/10/07 by Binbin Zhang, Hang Lv, Zhang, Binbin +21 · 30 citations
Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing
- M2MeT: The ICASSP 2022 Multi-Channel Multi-Party Meeting Transcription Challenge
2021/10/14 by Fan Yu, Yu, Fan, Shiliang Zhang +21 · 17 citations
Computer Science · #Speech Recognition and Synthesis #Music and Audio Processing #Topic Modeling
- Recent Developments on ESPnet Toolkit Boosted by Conformer
2020/10/26 by Guo, Pengcheng, Boyer, Florian, Chang, Xuankai +12 · 5 citations
#Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
- Exploring Speech Recognition, Translation, and Understanding with Discrete Speech Units: A Comparative Study
2023/09/27 by Xuankai Chang, Chang, Xuankai, Brian Yan +31 · 7 citations
Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and dialogue systems #electronic engineering #information engineering
- Summary On The ICASSP 2022 Multi-Channel Multi-Party Meeting Transcription Grand Challenge
2022/02/08 by Fan Yu, Yu, Fan, Shiliang Zhang +29 · 4 citations
Computer Science · #Speech Recognition and Synthesis #Music and Audio Processing #Speech and Audio Processing
- Unleashing the Power of Data Tsunami: A Comprehensive Survey on Data Assessment and Selection for Instruction Tuning of Language Models
2024/08/04 by Yulei Qin, Yuncheng Yang, Qin, Yulei +17 · 7 citations
Computer Science · #Natural Language Processing Techniques
- Distinctive and Natural Speaker Anonymization via Singular Value Transformation-assisted Matrix
2024/05/17 by Jixun Yao, Yao, Jixun, Qing Wang +7 · 6 citations
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Electrical engineering #Natural Language Processing Techniques #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- Contextualized End-to-End Speech Recognition with Contextual Phrase Prediction Network
2023/05/21 by Huang, Kaixun, Zhang, Ao, Yang, Zhanheng +4 · 4 citations
#Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
- OSUM: Advancing Open Speech Understanding Models with Limited Resources in Academia
2025/01/23 by Geng, Xuelong, Qijie Shao, Wei, Kun +34 · 11 citations
Computer Science · #Natural Language Processing Techniques
- Distinguishable Speaker Anonymization based on Formant and Fundamental Frequency Scaling
2022/11/06 by Yao, Jixun, Wang, Qing, Lei, Yi +4 · 3 citations
#Audio and Speech Processing (eess.AS) #Cryptography and Security (cs.CR) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
- Boundary and Context Aware Training for CIF-based Non-Autoregressive End-to-end ASR
2021/04/10 by Fan Yu, Haoneng Luo, Yu, Fan +15 · 3 citations
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- Adaptive Contextual Biasing for Transducer Based Streaming Speech Recognition
2023/06/01 by Xu, Tianyi, Yang, Zhanheng, Huang, Kaixun +6 · 3 citations
#Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
- Unveiling the Potential of LLM-Based ASR on Chinese Open-Source Datasets
2024/05/03 by Xuelong Geng, Geng, Xuelong, Tianyi Xu +21 · 6 citations
Computer Science · #Natural Language Processing Techniques
- Preserving background sound in noise-robust voice conversion via multi-task learning
2022/11/06 by Yao, Jixun, Lei, Yi, Wang, Qing +6 · 3 citations
#Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
- MUSA: Multi-lingual Speaker Anonymization via Serial Disentanglement
2024/07/16 by Yao, Jixun, Wang, Qing, Guo, Pengcheng +4 · 4 citations
#Audio and Speech Processing (eess.AS) #FOS: Electrical engineering #electronic engineering #information engineering
- Automatic channel selection and spatial feature integration for multi-channel speech recognition across various array topologies
2023/12/15 by Mu, Bingshen, Guo, Pengcheng, Guo, Dake +3 · 3 citations
#Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
- NPU-NTU System for Voice Privacy 2024 Challenge
2024/09/06 by Yao, Jixun, Kuzmin, Nikita, Wang, Qing +6 · 3 citations
#Audio and Speech Processing (eess.AS) #FOS: Electrical engineering #electronic engineering #information engineering
- Inaudible Adversarial Perturbations for Targeted Attack in Speaker Recognition
2020/05/21 by Qing Wang, Pengcheng Guo, Wang, Qing +3 · 1 citation
Computer Science · #Adversarial Robustness in Machine Learning #Anomaly Detection Techniques and Applications #Audio and Speech Processing (eess.AS) #Digital Media Forensic Detection #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
- The 2020 ESPnet update: new features, broadened applications, performance improvements, and future plans
2020/12/23 by Shinji Watanabe, Florian Boyer, Watanabe, Shinji +27 · 1 citation
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- SQ-Whisper: Speaker-Querying based Whisper Model for Target-Speaker ASR
2024/12/07 by Pengcheng Guo, Xuankai Chang, Guo, Pengcheng +7 · 3 citations
Computer Science · #Speech Recognition and Synthesis #Speech and dialogue systems
- Multi-Speaker ASR Combining Non-Autoregressive Conformer CTC and Conditional Speaker Chain
2021/06/16 by Guo, Pengcheng, Chang, Xuankai, Watanabe, Shinji +1 · 1 citation
#Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
- DiffAttack: Diffusion-based Timbre-reserved Adversarial Attack in Speaker Identification
2025/01/09 by Wang, Qing, Yao, Jixun, Sun, Zhaokai +3 · 3 citations
#Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
- Improving Transformer-based Conversational ASR by Inter-Sentential Attention Mechanism
2022/07/02 by Wei, Kun, Guo, Pengcheng, Jiang, Ning · 1 citation
#Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
- MFCCA:Multi-Frame Cross-Channel attention for multi-speaker ASR in Multi-party meeting scenario
2022/10/11 by Yu, Fan, Zhang, Shiliang, Guo, Pengcheng +4 · 1 citation
#Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
- The NPU-ASLP System for Audio-Visual Speech Recognition in MISP 2022 Challenge
2023/03/11 by Pengcheng Guo, Guo, Pengcheng, He Wang +7 · 1 citation
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Electrical engineering #Music and Audio Processing #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- Linguistic-Acoustic Similarity Based Accent Shift for Accent Recognition
2022/04/07 by Qijie Shao, Shao, Qijie, Jinghao Yan +11 · 1 citation
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- NWPU-ASLP System for the VoicePrivacy 2022 Challenge
2022/09/24 by Yao, Jixun, Wang, Qing, Zhang, Li +3 · 1 citation
#Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
- ICMC-ASR: The ICASSP 2024 In-Car Multi-Channel Automatic Speech Recognition Challenge
2024/01/07 by Wang, He, Guo, Pengcheng, Li, Yue +13 · 1 citation
#Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
- An audio-quality-based multi-strategy approach for target speaker extraction in the MISP 2023 Challenge
2024/01/08 by Han, Runduo, Yan, Xiaopeng, Xu, Weiming +6 · 1 citation
#Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
- Leveraging Open Knowledge for Advancing Task Expertise in Large Language Models
2024/08/28 by Yang, Yuncheng, Qin, Yulei, Wu, Tong +9 · 1 citation
#Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences
- Towards Rehearsal-Free Multilingual ASR: A LoRA-based Case Study on Whisper
2024/08/20 by Tianyi Xu, Kaixun Huang, Xu, Tianyi +11 · 1 citation
Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Hate Speech and Cyberbullying Detection #Sound (cs.SD) #electronic engineering #information engineering
- Mixture of LoRA Experts with Multi-Modal and Multi-Granularity LLM Generative Error Correction for Accented Speech Recognition
2025/07/12 by Mu, Bingshen, Wei, Kun, Guo, Pengcheng +1 · 4 citations
#Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
- TVDO: Tchebycheff Value-Decomposition Optimization for Multi-Agent Reinforcement Learning
2023/06/24 by Hu, Xiaoliang, Guo, Pengcheng, Li, Yadong +3 · 1 citation
#FOS: Computer and information sciences #Multiagent Systems (cs.MA)