vix.ing · top · new · best · stats · spec

Guo, Pengcheng

  1. WenetSpeech: A 10000+ Hours Multi-domain Mandarin Corpus for Speech Recognition
    2021/10/07 by Binbin Zhang, Hang Lv, Zhang, Binbin +21 · 30 citations
    Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing
  2. M2MeT: The ICASSP 2022 Multi-Channel Multi-Party Meeting Transcription Challenge
    2021/10/14 by Fan Yu, Yu, Fan, Shiliang Zhang +21 · 17 citations
    Computer Science · #Speech Recognition and Synthesis #Music and Audio Processing #Topic Modeling
  3. Recent Developments on ESPnet Toolkit Boosted by Conformer
    2020/10/26 by Guo, Pengcheng, Boyer, Florian, Chang, Xuankai +12 · 5 citations
    #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
  4. Exploring Speech Recognition, Translation, and Understanding with Discrete Speech Units: A Comparative Study
    2023/09/27 by Xuankai Chang, Chang, Xuankai, Brian Yan +31 · 7 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and dialogue systems #electronic engineering #information engineering
  5. Summary On The ICASSP 2022 Multi-Channel Multi-Party Meeting Transcription Grand Challenge
    2022/02/08 by Fan Yu, Yu, Fan, Shiliang Zhang +29 · 4 citations
    Computer Science · #Speech Recognition and Synthesis #Music and Audio Processing #Speech and Audio Processing
  6. Unleashing the Power of Data Tsunami: A Comprehensive Survey on Data Assessment and Selection for Instruction Tuning of Language Models
    2024/08/04 by Yulei Qin, Yuncheng Yang, Qin, Yulei +17 · 7 citations
    Computer Science · #Natural Language Processing Techniques
  7. Distinctive and Natural Speaker Anonymization via Singular Value Transformation-assisted Matrix
    2024/05/17 by Jixun Yao, Yao, Jixun, Qing Wang +7 · 6 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Electrical engineering #Natural Language Processing Techniques #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  8. Contextualized End-to-End Speech Recognition with Contextual Phrase Prediction Network
    2023/05/21 by Huang, Kaixun, Zhang, Ao, Yang, Zhanheng +4 · 4 citations
    #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
  9. OSUM: Advancing Open Speech Understanding Models with Limited Resources in Academia
    2025/01/23 by Geng, Xuelong, Qijie Shao, Wei, Kun +34 · 11 citations
    Computer Science · #Natural Language Processing Techniques
  10. Distinguishable Speaker Anonymization based on Formant and Fundamental Frequency Scaling
    2022/11/06 by Yao, Jixun, Wang, Qing, Lei, Yi +4 · 3 citations
    #Audio and Speech Processing (eess.AS) #Cryptography and Security (cs.CR) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
  11. Boundary and Context Aware Training for CIF-based Non-Autoregressive End-to-end ASR
    2021/04/10 by Fan Yu, Haoneng Luo, Yu, Fan +15 · 3 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  12. Adaptive Contextual Biasing for Transducer Based Streaming Speech Recognition
    2023/06/01 by Xu, Tianyi, Yang, Zhanheng, Huang, Kaixun +6 · 3 citations
    #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
  13. Unveiling the Potential of LLM-Based ASR on Chinese Open-Source Datasets
    2024/05/03 by Xuelong Geng, Geng, Xuelong, Tianyi Xu +21 · 6 citations
    Computer Science · #Natural Language Processing Techniques
  14. Preserving background sound in noise-robust voice conversion via multi-task learning
    2022/11/06 by Yao, Jixun, Lei, Yi, Wang, Qing +6 · 3 citations
    #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
  15. MUSA: Multi-lingual Speaker Anonymization via Serial Disentanglement
    2024/07/16 by Yao, Jixun, Wang, Qing, Guo, Pengcheng +4 · 4 citations
    #Audio and Speech Processing (eess.AS) #FOS: Electrical engineering #electronic engineering #information engineering
  16. Automatic channel selection and spatial feature integration for multi-channel speech recognition across various array topologies
    2023/12/15 by Mu, Bingshen, Guo, Pengcheng, Guo, Dake +3 · 3 citations
    #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
  17. NPU-NTU System for Voice Privacy 2024 Challenge
    2024/09/06 by Yao, Jixun, Kuzmin, Nikita, Wang, Qing +6 · 3 citations
    #Audio and Speech Processing (eess.AS) #FOS: Electrical engineering #electronic engineering #information engineering
  18. Inaudible Adversarial Perturbations for Targeted Attack in Speaker Recognition
    2020/05/21 by Qing Wang, Pengcheng Guo, Wang, Qing +3 · 1 citation
    Computer Science · #Adversarial Robustness in Machine Learning #Anomaly Detection Techniques and Applications #Audio and Speech Processing (eess.AS) #Digital Media Forensic Detection #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
  19. The 2020 ESPnet update: new features, broadened applications, performance improvements, and future plans
    2020/12/23 by Shinji Watanabe, Florian Boyer, Watanabe, Shinji +27 · 1 citation
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  20. SQ-Whisper: Speaker-Querying based Whisper Model for Target-Speaker ASR
    2024/12/07 by Pengcheng Guo, Xuankai Chang, Guo, Pengcheng +7 · 3 citations
    Computer Science · #Speech Recognition and Synthesis #Speech and dialogue systems
  21. Multi-Speaker ASR Combining Non-Autoregressive Conformer CTC and Conditional Speaker Chain
    2021/06/16 by Guo, Pengcheng, Chang, Xuankai, Watanabe, Shinji +1 · 1 citation
    #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
  22. DiffAttack: Diffusion-based Timbre-reserved Adversarial Attack in Speaker Identification
    2025/01/09 by Wang, Qing, Yao, Jixun, Sun, Zhaokai +3 · 3 citations
    #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
  23. Improving Transformer-based Conversational ASR by Inter-Sentential Attention Mechanism
    2022/07/02 by Wei, Kun, Guo, Pengcheng, Jiang, Ning · 1 citation
    #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
  24. MFCCA:Multi-Frame Cross-Channel attention for multi-speaker ASR in Multi-party meeting scenario
    2022/10/11 by Yu, Fan, Zhang, Shiliang, Guo, Pengcheng +4 · 1 citation
    #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
  25. The NPU-ASLP System for Audio-Visual Speech Recognition in MISP 2022 Challenge
    2023/03/11 by Pengcheng Guo, Guo, Pengcheng, He Wang +7 · 1 citation
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Electrical engineering #Music and Audio Processing #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  26. Linguistic-Acoustic Similarity Based Accent Shift for Accent Recognition
    2022/04/07 by Qijie Shao, Shao, Qijie, Jinghao Yan +11 · 1 citation
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  27. NWPU-ASLP System for the VoicePrivacy 2022 Challenge
    2022/09/24 by Yao, Jixun, Wang, Qing, Zhang, Li +3 · 1 citation
    #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
  28. ICMC-ASR: The ICASSP 2024 In-Car Multi-Channel Automatic Speech Recognition Challenge
    2024/01/07 by Wang, He, Guo, Pengcheng, Li, Yue +13 · 1 citation
    #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
  29. An audio-quality-based multi-strategy approach for target speaker extraction in the MISP 2023 Challenge
    2024/01/08 by Han, Runduo, Yan, Xiaopeng, Xu, Weiming +6 · 1 citation
    #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
  30. Leveraging Open Knowledge for Advancing Task Expertise in Large Language Models
    2024/08/28 by Yang, Yuncheng, Qin, Yulei, Wu, Tong +9 · 1 citation
    #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences
  31. Towards Rehearsal-Free Multilingual ASR: A LoRA-based Case Study on Whisper
    2024/08/20 by Tianyi Xu, Kaixun Huang, Xu, Tianyi +11 · 1 citation
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Hate Speech and Cyberbullying Detection #Sound (cs.SD) #electronic engineering #information engineering
  32. Mixture of LoRA Experts with Multi-Modal and Multi-Granularity LLM Generative Error Correction for Accented Speech Recognition
    2025/07/12 by Mu, Bingshen, Wei, Kun, Guo, Pengcheng +1 · 4 citations
    #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
  33. TVDO: Tchebycheff Value-Decomposition Optimization for Multi-Agent Reinforcement Learning
    2023/06/24 by Hu, Xiaoliang, Guo, Pengcheng, Li, Yadong +3 · 1 citation
    #FOS: Computer and information sciences #Multiagent Systems (cs.MA)