vix.ing · top · new · best · stats · spec

Pengyuan Zhang

  1. Emilia: An Extensive, Multilingual, and Diverse Speech Dataset for Large-Scale Speech Generation
    2024/07/07 by Haorui He, Zengqiang Shang, He, Haorui +25 · 124 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Speech Recognition and Synthesis #Speech and dialogue systems #Topic Modeling #electronic engineering #information engineering
  2. DPT-FSNet: Dual-path Transformer Based Full-band and Sub-band Fusion Network for Speech Enhancement
    2021/04/27 by Dang Feng, Hangting Chen, Dang, Feng +3 · 17 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  3. Open Source MagicData-RAMC: A Rich Annotated Mandarin Conversational(RAMC) Speech Dataset
    2022/03/31 by Zehui Yang, Yifan Chen, Yang, Zehui +21 · 18 citations
    Computer Science · #Speech Recognition and Synthesis #Speech and dialogue systems #Topic Modeling
  4. CN-CELEB: a challenging Chinese speaker recognition dataset
    2019/10/31 by Yue Fan, Fan, Yue, Jiawen Kang +17 · 15 citations
    Computer Science · #Speech Recognition and Synthesis #Music and Audio Processing #Speech and Audio Processing
  5. Improving Short Utterance Anti-Spoofing with AASIST2
    2023/09/15 by Yuxiang Zhang, Jingze Lu, Zhang, Yuxiang +7 · 8 citations
    Computer Science · Medicine · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #Voice and Speech Disorders #electronic engineering #information engineering
  6. Improving non-autoregressive end-to-end speech recognition with pre-trained acoustic and language models
    2022/01/25 by Keqi Deng, Zehui Yang, Deng, Keqi +9 · 4 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  7. Boosting Cross-Domain Speech Recognition with Self-Supervision
    2022/06/20 by Han Zhu, Gaofeng Cheng, Zhu, Han +9 · 4 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  8. The Impact of Silence on Speech Anti-Spoofing
    2023/09/21 by Yuxiang Zhang, Zhuo Li, Zhang, Yuxiang +9 · 4 citations
    Computer Science · Medicine · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #Voice and Speech Disorders #electronic engineering #information engineering
  9. Emilia: A Large-Scale, Extensive, Multilingual, and Diverse Dataset for Speech Generation
    2025/01/27 by Haorui He, He, Haorui, Zengqiang Shang +25 · 12 citations
    Computer Science · #Speech Recognition and Synthesis #Speech and dialogue systems
  10. PCF: ECAPA-TDNN with Progressive Channel Fusion for Speaker Verification
    2023/03/01 by Zhenduo Zhao, Zhao, Zhenduo, Zhuo Li +5 · 3 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  11. Improving CTC-based speech recognition via knowledge transferring from pre-trained language models
    2022/02/22 by Keqi Deng, Songjun Cao, Deng, Keqi +11 · 2 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Topic Modeling #electronic engineering #information engineering
  12. Alternative Pseudo-Labeling for Semi-Supervised Automatic Speech Recognition
    2023/08/12 by Zhu Han, Dongji Gao, Zhu, Han +9 · 2 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  13. SLIDE: Integrating Speech Language Model with LLM for Spontaneous Spoken Dialogue Generation
    2025/01/01 by Haitian Lu, Lu, Haitian, Gaofeng Cheng +9 · 4 citations
    Computer Science · #Speech and dialogue systems #Natural Language Processing Techniques #Topic Modeling
  14. Transformer-based Online CTC/attention End-to-End Speech Recognition Architecture
    2020/01/15 by Haoran Miao, Gaofeng Cheng, Miao, Haoran +7 · 1 citation
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Music and Audio Processing #Natural Language Processing Techniques #Neural and Evolutionary Computing (cs.NE) #Sound (cs.SD) #Speech Recognition and Synthesis #electronic engineering #information engineering
  15. Improved Conformer-based End-to-End Speech Recognition Using Neural Architecture Search
    2021/04/12 by Yukun Liu, Liu, Yukun, Li Ta +5 · 1 citation
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Topic Modeling #electronic engineering #information engineering
  16. Wav2vec-S: Semi-Supervised Pre-Training for Low-Resource ASR
    2021/10/09 by Han Zhu, Li Wang, Zhu, Han +9 · 1 citation
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #electronic engineering #information engineering
  17. Decoupled Federated Learning for ASR with Non-IID Data
    2022/06/18 by Han Zhu, Zhu, Han, Jindong Wang +7 · 1 citation
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #Distributed #FOS: Computer and information sciences #FOS: Electrical engineering #Internet Traffic Analysis and Secure E-voting #Parallel #Privacy-Preserving Technologies in Data #Sound (cs.SD) #Speech Recognition and Synthesis #and Cluster Computing (cs.DC) #electronic engineering #information engineering
  18. Streaming non-autoregressive model for any-to-many voice conversion
    2022/06/15 by Ziyi Chen, Haoran Miao, Chen, Ziyi +3 · 1 citation
    Computer Science · #Speech Recognition and Synthesis #Speech and Audio Processing #Advanced Data Compression Techniques
  19. Synthetic Speech Detection Based on Temporal Consistency and Distribution of Speaker Features
    2023/09/29 by Yuxiang Zhang, Zhuo Li, Zhang, Yuxiang +7 · 1 citation
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  20. SF-Speech: Straightened Flow for Zero-Shot Voice Clone
    2024/10/16 by Xuyuan Li, Zengqiang Shang, Li, Xuyuan +11 · 2 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #Speech Recognition and Synthesis #electronic engineering #information engineering
  21. DSNet: Disentangled Siamese Network with Neutral Calibration for Speech Emotion Recognition
    2023/12/25 by Chengxin Chen, Chen, Chengxin, Pengyuan Zhang +1 · 1 citation
    Computer Science · Psychology · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Emotion and Mood Recognition #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  22. Modality-Collaborative Transformer with Hybrid Feature Reconstruction for Robust Emotion Recognition
    2023/12/26 by Chengxin Chen, Pengyuan Zhang, Chen, Chengxin +1 · 1 citation
    Computer Science · Psychology · Social Sciences · #Advanced Computing and Algorithms #Artificial Intelligence (cs.AI) #Computer Vision and Pattern Recognition (cs.CV) #Emotion and Mood Recognition #FOS: Computer and information sciences #Sentiment Analysis and Opinion Mining