vix.ing · top · new · best · stats · spec

Liu, Xunying

  1. WavLLM: Towards Robust and Adaptive Speech Large Language Model
    2024/03/31 by Shujie Hu, Long Zhou, Hu, Shujie +20 · 54 citations
    Computer Science · #Natural Language Processing Techniques #Speech Recognition and Synthesis #Topic Modeling
  2. VQMIVC: Vector Quantization and Mutual Information-Based Unsupervised Speech Representation Disentanglement for One-shot Voice Conversion
    2021/06/18 by Disong Wang, Wang, Disong, Liqun Deng +9 · 14 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Multimedia (cs.MM) #Music and Audio Processing #Signal Processing (eess.SP) #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  3. GigaSpeech 2: An Evolving, Large-Scale and Multi-domain ASR Corpus for Low-Resource Languages with Automated Crawling, Transcription and Refinement
    2024/06/17 by Yifan Yang, Yang, Yifan, Zheshu Song +29 · 18 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Sound (cs.SD) #electronic engineering #information engineering
  4. Replay and Synthetic Speech Detection with Res2net Architecture
    2020/10/28 by Xu Li, Na Li, Li, Xu +11 · 6 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  5. Adversarial Attacks on GMM i-vector based Speaker Verification Systems
    2019/11/08 by Li, Xu, Zhong, Jinghua, Wu, Xixin +3 · 5 citations
    #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #Cryptography and Security (cs.CR) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Signal Processing (eess.SP) #electronic engineering #information engineering
  6. Large Language Model Can Transcribe Speech in Multi-Talker Scenarios with Versatile Instructions
    2024/09/13 by Lingwei Meng, Meng, Lingwei, Shujie Hu +15 · 14 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #Speech Recognition and Synthesis #electronic engineering #information engineering
  7. Empowering Whisper as a Joint Multi-Talker and Target-Talker Speech Recognition System
    2024/07/13 by Lingwei Meng, Jiawen Kang, Meng, Lingwei +11 · 11 citations
    Computer Science · Environmental Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #Educational Reforms and Innovations #FOS: Computer and information sciences #FOS: Electrical engineering #Hate Speech and Cyberbullying Detection #Sound (cs.SD) #electronic engineering #information engineering
  8. Enhancing Pre-trained ASR System Fine-tuning for Dysarthric Speech Recognition using Adversarial Data Augmentation
    2024/01/01 by Wang, Huimeng, Jin, Zengrui, Geng, Mengzhe +5 · 9 citations
    #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
  9. Channel-wise Gated Res2Net: Towards Robust Detection of Synthetic Speech Attacks
    2021/07/19 by Xu Li, Li, Xu, Xixin Wu +7 · 4 citations
    Computer Science · #Speech Recognition and Synthesis #Speech and Audio Processing #Music and Audio Processing
  10. Learning Explicit Prosody Models and Deep Speaker Embeddings for Atypical Voice Conversion
    2020/11/03 by Disong Wang, Songxiang Liu, Wang, Disong +9 · 3 citations
    Computer Science · Medicine · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #Voice and Speech Disorders #electronic engineering #information engineering
  11. Investigating Robustness of Adversarial Samples Detection for Automatic Speaker Verification
    2020/06/11 by Xu Li, Li, Xu, Na Li +13 · 3 citations
    Computer Science · #Adversarial Robustness in Machine Learning #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Sound (cs.SD) #Speech Recognition and Synthesis #Wireless Signal Modulation Classification #electronic engineering #information engineering
  12. Audio-visual Recognition of Overlapped speech for the LRS2 dataset
    2020/01/06 by Jianwei Yu, Shi-Xiong Zhang, Yu, Jianwei +17 · 3 citations
    Computer Science · #Speech and Audio Processing #Music and Audio Processing #Speech Recognition and Synthesis
  13. Exploring linguistic feature and model combination for speech recognition based automatic AD detection
    2022/06/28 by Wang, Yi, Wang, Tianzi, Ye, Zi +5 · 3 citations
    #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #electronic engineering #information engineering
  14. Conformer Based Elderly Speech Recognition System for Alzheimer's Disease Detection
    2022/06/23 by Tianzi Wang, Jiajun Deng, Wang, Tianzi +17 · 4 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Sound (cs.SD) #Speech Recognition and Synthesis #electronic engineering #information engineering
  15. DSNAS: Direct Neural Architecture Search without Parameter Retraining
    2020/02/21 by Shoukang Hu, Sirui Xie, Hu, Shoukang +11 · 2 citations
    Computer Science · #Advanced Neural Network Applications #Domain Adaptation and Few-Shot Learning #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Machine Learning and ELM
  16. Self-supervised ASR Models and Features For Dysarthric and Elderly Speech Recognition
    2024/07/03 by Shujie Hu, Xurong Xie, Hu, Shujie +19 · 9 citations
    Computer Science · Medicine · #Speech Recognition and Synthesis #Voice and Speech Disorders
  17. Leveraging Pretrained Representations with Task-related Keywords for Alzheimer's Disease Detection
    2023/03/14 by Jinchao Li, Kaitao Song, Li, Jinchao +13 · 3 citations
    Psychology · Computer Science · #Emotion and Mood Recognition #Speech Recognition and Synthesis
  18. Bayesian Transformer Language Models for Speech Recognition
    2021/02/09 by Boyang Xue, Xue, Boyang, Jianwei Yu +15 · 2 citations
    Computer Science · #Speech Recognition and Synthesis #Topic Modeling #Natural Language Processing Techniques
  19. Speaker Adaptation Using Spectro-Temporal Deep Features for Dysarthric and Elderly Speech Recognition
    2022/02/21 by Mengzhe Geng, Geng, Mengzhe, Xurong Xie +13 · 5 citations
    Health Professions · Medicine · Psychology · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Dysphagia Assessment and Management #FOS: Biological sciences #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Phonetics and Phonology Research #Quantitative Methods (q-bio.QM) #Sound (cs.SD) #Voice and Speech Disorders #electronic engineering #information engineering
  20. Personalized Adversarial Data Augmentation for Dysarthric and Elderly Speech Recognition
    2022/05/13 by Zengrui Jin, Mengzhe Geng, Jin, Zengrui +11 · 3 citations
    Computer Science · Medicine · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Speech Recognition and Synthesis #Speech and Audio Processing #Voice and Speech Disorders #electronic engineering #information engineering
  21. On-the-Fly Feature Based Rapid Speaker Adaptation for Dysarthric and Elderly Speech Recognition
    2022/03/28 by Geng, Mengzhe, Xie, Xurong, Su, Rongfeng +7 · 2 citations
    #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Sound (cs.SD) #electronic engineering #information engineering
  22. Mixed Precision of Quantization of Transformer Language Models for Speech Recognition
    2021/11/29 by Xu, Junhao, Hu, Shoukang, Yu, Jianwei +2 · 2 citations
    #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
  23. Disentangled Speech Representation Learning for One-Shot Cross-lingual Voice Conversion Using β-VAE
    2022/10/25 by Hui Lü, Lu, Hui, Disong Wang +9 · 2 citations
    Computer Science · Medicine · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #Voice and Speech Disorders #electronic engineering #information engineering
  24. Perceiver-Prompt: Flexible Speaker Adaptation in Whisper for Chinese Disordered Speech Recognition
    2024/06/14 by Yicong Jiang, Jiang, Yicong, Tianzi Wang +17 · 4 citations
    Computer Science · Environmental Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Educational Reforms and Innovations #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #Speech Recognition and Synthesis #electronic engineering #information engineering
  25. Audio-visual End-to-end Multi-channel Speech Separation, Dereverberation and Recognition
    2023/07/06 by Guinan Li, Jiajun Deng, Li, Guinan +15 · 2 citations
    Computer Science · Engineering · #Speech and Audio Processing #Advanced Adaptive Filtering Techniques #Blind Source Separation Techniques
  26. Confidence Score Based Conformer Speaker Adaptation for Speech Recognition
    2022/06/24 by Jiajun Deng, Deng, Jiajun, Xurong Xie +17 · 2 citations
    Computer Science · #Advanced Data Compression Techniques #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  27. Towards Automatic Data Augmentation for Disordered Speech Recognition
    2023/12/14 by Jin, Zengrui, Xie, Xurong, Wang, Tianzi +5 · 2 citations
    #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
  28. Audio-visual Multi-channel Recognition of Overlapped Speech
    2020/05/18 by Yu, Jianwei, Wu, Bo, Gu, Rongzhi +7 · 1 citation
    #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
  29. Bayesian x-vector: Bayesian Neural Network based x-vector System for Speaker Verification
    2020/04/08 by Xu Li, Jinghua Zhong, Li, Xu +11 · 1 citation
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  30. Transferring Source Style in Non-Parallel Voice Conversion
    2020/05/19 by Songxiang Liu, Liu, Songxiang, Yuewen Cao +13 · 1 citation
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Electrical engineering #Speech Recognition and Synthesis #Speech and Audio Processing #Speech and dialogue systems #electronic engineering #information engineering
  31. Improved End-to-End Dysarthric Speech Recognition via Meta-learning Based Model Re-initialization
    2020/11/03 by Wang, Disong, Yu, Jianwei, Wu, Xixin +3 · 1 citation
    #Audio and Speech Processing (eess.AS) #FOS: Electrical engineering #electronic engineering #information engineering
  32. Audio-visual Multi-channel Integration and Recognition of Overlapped Speech
    2020/11/16 by Jianwei Yu, Shi-Xiong Zhang, Yu, Jianwei +15 · 1 citation
    Computer Science · Engineering · #Advanced Adaptive Filtering Techniques #Audio and Speech Processing (eess.AS) #FOS: Electrical engineering #Music and Audio Processing #Speech and Audio Processing #electronic engineering #information engineering
  33. Disentangling Speakers in Multi-Talker Speech Recognition with Speaker-Aware CTC
    2024/09/19 by Jiawen Kang, Lingwei Meng, Kang, Jiawen +11 · 3 citations
    Computer Science · #Advanced Data Compression Techniques #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  34. Unsupervised Domain Adaptation for Dysarthric Speech Detection via Domain Adversarial Training and Mutual Information Minimization
    2021/06/18 by Disong Wang, Wang, Disong, Liqun Deng +8 · 1 citation
    Computer Science · Medicine · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Signal Processing (eess.SP) #Sound (cs.SD) #Speech Recognition and Synthesis #Voice and Speech Disorders #electronic engineering #information engineering
  35. VAENAR-TTS: Variational Auto-Encoder based Non-AutoRegressive Text-to-Speech Synthesis
    2021/07/07 by Hui Lu, Zhiyong Wu, Lu, Hui +11 · 1 citation
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Multimedia (cs.MM) #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  36. Adversarial Data Augmentation for Disordered Speech Recognition
    2021/08/02 by Zengrui Jin, Mengzhe Geng, Jin, Zengrui +11 · 1 citation
    Computer Science · Medicine · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #Voice and Speech Disorders #electronic engineering #information engineering
  37. Mixed Precision DNN Qunatization for Overlapped Speech Separation and Recognition
    2021/11/29 by Junhao Xu, Jianwei Yu, Xu, Junhao +5 · 1 citation
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  38. Exploring SSL Discrete Tokens for Multilingual ASR
    2024/09/13 by Mingyu Cui, Cui, Mingyu, Daxin Tan +12 · 2 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and dialogue systems #electronic engineering #information engineering
  39. Use of Speech Impairment Severity for Dysarthric Speech Recognition
    2023/05/18 by Mengzhe Geng, Zengrui Jin, Geng, Mengzhe +17 · 2 citations
    Computer Science · Medicine · Psychology · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Phonetics and Phonology Research #Sound (cs.SD) #Speech Recognition and Synthesis #Voice and Speech Disorders #electronic engineering #information engineering
  40. Exploiting Audio-Visual Features with Pretrained AV-HuBERT for Multi-Modal Dysarthric Speech Reconstruction
    2024/01/31 by Xueyuan Chen, Chen, Xueyuan, Yuejiao Wang +11 · 2 citations
    Computer Science · Medicine · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #Voice and Speech Disorders #electronic engineering #information engineering
  41. Adversarial Data Augmentation Using VAE-GAN for Disordered Speech Recognition
    2022/11/03 by Jin, Zengrui, Xie, Xurong, Geng, Mengzhe +5 · 1 citation
    #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
  42. Bayesian Neural Network Language Modeling for Speech Recognition
    2022/08/28 by Boyang Xue, Xue, Boyang, Shoukang Hu +9 · 1 citation
    Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Neural Networks and Applications #Speech Recognition and Synthesis #Speech and Audio Processing
  43. Factorised Speaker-environment Adaptive Training of Conformer Speech Recognition Systems
    2023/06/26 by Jiajun Deng, Deng, Jiajun, Guinan Li +15 · 1 citation
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  44. Neural Architecture Search For LF-MMI Trained Time Delay Neural Networks
    2022/01/08 by Hu, Shoukang, Xie, Xurong, Cui, Mingyu +6 · 1 citation
    #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
  45. Bayesian Learning of LF-MMI Trained Time Delay Neural Networks for Speech Recognition
    2020/12/08 by Hu, Shoukang, Xie, Xurong, Liu, Shansong +5 · 1 citation
    #Audio and Speech Processing (eess.AS) #FOS: Electrical engineering #electronic engineering #information engineering
  46. Exploiting prompt learning with pre-trained language models for Alzheimer's Disease detection
    2022/10/29 by Yi Wang, Wang, Yi, Jiajun Deng +11 · 1 citation
    Health Professions · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Interpreting and Communication in Healthcare #Machine Learning (cs.LG) #Sound (cs.SD) #electronic engineering #information engineering
  47. Exploring Self-supervised Pre-trained ASR Models For Dysarthric and Elderly Speech Recognition
    2023/02/28 by Shujie Hu, Hu, Shujie, Xurong Xie +14 · 2 citations
    Computer Science · Medicine · Psychology · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Phonetics and Phonology Research #Sound (cs.SD) #Speech Recognition and Synthesis #Voice and Speech Disorders #electronic engineering #information engineering
  48. Hyper-parameter Adaptation of Conformer ASR Systems for Elderly and Dysarthric Speech Recognition
    2023/06/27 by Tianzi Wang, Shoukang Hu, Wang, Tianzi +13 · 1 citation
    Computer Science · Medicine · Psychology · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Phonetics and Phonology Research #Speech Recognition and Synthesis #Voice and Speech Disorders #electronic engineering #information engineering
  49. Homogeneous Speaker Features for On-the-Fly Dysarthric and Elderly Speaker Adaptation
    2024/07/08 by Mengzhe Geng, Geng, Mengzhe, Xurong Xie +17 · 3 citations
    Medicine · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Human-Computer Interaction (cs.HC) #Machine Learning (cs.LG) #Sound (cs.SD) #Voice and Speech Disorders #electronic engineering #information engineering
  50. Towards Effective and Efficient Non-autoregressive Decoding Using Block-based Attention Mask
    2024/06/14 by Tianzi Wang, Wang, Tianzi, Xurong Xie +21 · 2 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Blind Source Separation Techniques #Digital Media Forensic Detection #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #Speech and Audio Processing #electronic engineering #information engineering
  51. Exploiting Cross-domain And Cross-Lingual Ultrasound Tongue Imaging Features For Elderly And Dysarthric Speech Recognition
    2022/06/15 by Shujie Hu, Hu, Shujie, Xurong Xie +14 · 2 citations
    Computer Science · Medicine · Psychology · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Phonetics and Phonology Research #Speech Recognition and Synthesis #Voice and Speech Disorders #electronic engineering #information engineering
  52. One-pass Multiple Conformer and Foundation Speech Systems Compression and Quantization Using An All-in-one Neural Model
    2024/06/14 by Zhaoqing Li, Li, Zhaoqing, Haoning Xu +17 · 3 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Neural Networks and Applications #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  53. Towards Effective and Compact Contextual Representation for Conformer Transducer Speech Recognition Systems
    2023/06/23 by Mingyu Cui, Cui, Mingyu, Jiawen Kang +11 · 1 citation
    Computer Science · Psychology · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Phonetics and Phonology Research #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  54. Exploiting Cross Domain Acoustic-to-articulatory Inverted Features For Disordered Speech Recognition
    2022/03/19 by Shujie Hu, Shansong Liu, Hu, Shujie +15 · 1 citation
    Computer Science · Medicine · Psychology · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Phonetics and Phonology Research #Speech Recognition and Synthesis #Voice and Speech Disorders #electronic engineering #information engineering