vix.ing · top · new · best · stats · spec

Lei Xie

  1. Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens
    2025/03/03 by Xinsheng Wang, Wang, Xinsheng, Mingqi Jiang +49 · 2 voices · 59 citations
    Computer Science · #Natural Language Processing Techniques #Speech Recognition and Synthesis #Speech and dialogue systems #cs.AI #cs.SD #eess.AS
  2. DCCRN: Deep Complex Convolution Recurrent Network for Phase-Aware Speech Enhancement
    2020/08/01 by Yanxin Hu, Yun Liu, Hu, Yanxin +15 · 34 citations
    Computer Science · Neuroscience · #Speech and Audio Processing #Hearing Loss and Rehabilitation #Speech Recognition and Synthesis
  3. WenetSpeech: A 10000+ Hours Multi-domain Mandarin Corpus for Speech Recognition
    2021/10/07 by Binbin Zhang, Zhang, Binbin, Hang Lv +21 · 30 citations
    Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing
  4. Freeze-Omni: A Smart and Low Latency Speech-to-speech Dialogue Model with Frozen LLM
    2024/11/01 by Xiong Wang, Wang, Xiong, Yangze Li +13 · 47 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and dialogue systems #Topic Modeling #electronic engineering #information engineering
  5. M2MeT: The ICASSP 2022 Multi-Channel Multi-Party Meeting Transcription Challenge
    2021/10/14 by Fan Yu, Shiliang Zhang, Yu, Fan +21 · 17 citations
    Computer Science · #Speech Recognition and Synthesis #Music and Audio Processing #Topic Modeling
  6. AISHELL-4: An Open Source Dataset for Speech Enhancement, Separation,\n Recognition and Speaker Diarization in Conference Scenario
    2021/04/08 by Yihui Fu, Luyao Cheng, Fu, Yihui +22 · 16 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  7. MambaAD: Exploring State Space Models for Multi-class Unsupervised Anomaly Detection
    2024/04/09 by Haoyang He, Yuhu Bai, He, Haoyang +17 · 23 citations
    Computer Science · Engineering · #Anomaly Detection Techniques and Applications #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Network Security and Intrusion Detection #Smart Grid Security and Resilience
  8. WeNet: Production oriented Streaming and Non-streaming End-to-End Speech Recognition Toolkit
    2021/02/02 by Zhuoyuan Yao, Yao, Zhuoyuan, Di Wu +17 · 12 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #electronic engineering #information engineering
  9. Open Source MagicData-RAMC: A Rich Annotated Mandarin Conversational(RAMC) Speech Dataset
    2022/03/31 by Zehui Yang, Yifan Chen, Yang, Zehui +21 · 11 citations
    Computer Science · #Speech Recognition and Synthesis #Speech and dialogue systems #Topic Modeling
  10. SELM: Speech Enhancement Using Discrete Tokens and Language Models
    2023/12/15 by Ziqian Wang, Xinfa Zhu, Wang, Ziqian +11 · 16 citations
    Computer Science · Medicine · #Speech and Audio Processing #Speech Recognition and Synthesis #Voice and Speech Disorders
  11. Time Domain Audio Visual Speech Separation
    2019/04/07 by Jian Wu, Yong Xu, Wu, Jian +11 · 7 citations
    Computer Science · Engineering · #Advanced Adaptive Filtering Techniques #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech and Audio Processing #electronic engineering #information engineering
  12. WeNet 2.0: More Productive End-to-End Speech Recognition Toolkit
    2022/03/29 by Binbin Zhang, Di Wu, Zhang, Binbin +17 · 9 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Topic Modeling #electronic engineering #information engineering
  13. Multi-band MelGAN: Faster Waveform Generation for High-Quality Text-to-Speech
    2020/05/11 by Geng Yang, Shan Yang, Yang, Geng +9 · 7 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  14. WenetSpeech4TTS: A 12,800-hour Mandarin TTS Corpus for Large Speech Generation Model Benchmark
    2024/06/09 by Linhan Ma, Ma, Linhan, Dake Guo +16 · 15 citations
    Computer Science · Psychology · #Audio and Speech Processing (eess.AS) #FOS: Electrical engineering #Phonetics and Phonology Research #Speech Recognition and Synthesis #Speech and dialogue systems #electronic engineering #information engineering
  15. Single-Codec: Single-Codebook Speech Codec towards High-Performance Speech Generation
    2024/06/11 by Hanzhao Li, Liumeng Xue, Li, Hanzhao +15 · 14 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Electrical engineering #Speech Recognition and Synthesis #Speech and dialogue systems #electronic engineering #information engineering
  16. Collaborative Sensing in Perceptive Mobile Networks: Opportunities and Challenges
    2022/05/31 by Lei Xie, S. H. Song, Xie, Lei +5 · 7 citations
    Computer Science · Engineering · #Energy Efficient Wireless Sensor Networks #Energy Harvesting in Wireless Networks #FOS: Computer and information sciences #FOS: Electrical engineering #Indoor and Outdoor Localization Technologies #Information Theory (cs.IT) #Signal Processing (eess.SP) #electronic engineering #information engineering
  17. GenSE: Generative Speech Enhancement via Language Models using Hierarchical Modeling
    2025/02/05 by Jixun Yao, Hexin Liu, Yao, Jixun +9 · 14 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #Speech and dialogue systems #electronic engineering #information engineering
  18. TFGAN: Time and Frequency Domain Based Generative Adversarial Network for High-fidelity Speech Synthesis
    2020/11/24 by Qiao Tian, Yi Chen, Tian, Qiao +11 · 4 citations
    Computer Science · #Speech Recognition and Synthesis #Speech and Audio Processing #Music and Audio Processing
  19. A Comprehensive Library for Benchmarking Multi-class Visual Anomaly Detection
    2024/06/05 by Jiangning Zhang, Haoyang He, Zhang, Jiangning +17 · 8 citations
    Computer Science · Medicine · #Anomaly Detection Techniques and Applications #COVID-19 diagnosis using AI #Computer Vision and Pattern Recognition (cs.CV) #Currency Recognition and Detection #FOS: Computer and information sciences
  20. METTS: Multilingual Emotional Text-to-Speech by Cross-speaker and Cross-lingual Emotion Transfer
    2023/07/29 by Xinfa Zhu, Yi Lei, Zhu, Xinfa +11 · 6 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Electrical engineering #Speech Recognition and Synthesis #electronic engineering #information engineering
  21. DSPGAN: a GAN-based universal vocoder for high-fidelity TTS by time-frequency domain supervision from DSP
    2022/11/02 by Kun Song, Song, Kun, Yongmao Zhang +13 · 5 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  22. PromptStyle: Controllable Style Transfer for Text-to-Speech with Natural Language Descriptions
    2023/05/31 by Liu, Guanghou, Yongmao Zhang, Zhang, Yongmao +10 · 5 citations
    Computer Science · #Speech Recognition and Synthesis #Speech and Audio Processing #Music and Audio Processing
  23. Molecular Mechanics-Driven Graph Neural Network with Multiplex Graph for Molecular Structures
    2020/11/15 by Shuo Zhang, Yang Liu, Zhang, Shuo +3 · 4 citations
    Materials Science · Computer Science · Biochemistry, Genetics and Molecular Biology · #Machine Learning in Materials Science #Computational Drug Discovery Methods #Protein Structure and Dynamics
  24. Summary On The ICASSP 2022 Multi-Channel Multi-Party Meeting Transcription Grand Challenge
    2022/02/08 by Fan Yu, Yu, Fan, Shiliang Zhang +29 · 4 citations
    Computer Science · #Speech Recognition and Synthesis #Music and Audio Processing #Speech and Audio Processing
  25. Conversational End-to-End TTS for Voice Agent
    2020/05/21 by Haohan Guo, Guo, Haohan, Shaofei Zhang +7 · 3 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Topic Modeling #electronic engineering #information engineering
  26. Perceptive Mobile Network with Distributed Target Monitoring Terminals: Leaking Communication Energy for Sensing
    2021/12/29 by Lei Xie, Xie, Lei, Peilan Wang +5 · 4 citations
    Engineering · #Energy Harvesting in Wireless Networks #FOS: Computer and information sciences #FOS: Electrical engineering #Full-Duplex Wireless Communications #Indoor and Outdoor Localization Technologies #Information Theory (cs.IT) #Signal Processing (eess.SP) #electronic engineering #information engineering
  27. Uformer: A Unet based dilated complex & real dual-path conformer network for simultaneous speech enhancement and dereverberation
    2021/11/11 by Yihui Fu, Fu, Yihui, Yun Liu +11 · 4 citations
    Computer Science · Health Professions · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Infant Health and Development #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  28. MMGER: Multi-modal and Multi-granularity Generative Error Correction with LLM for Joint Accent and Speech Recognition
    2024/05/06 by Bingshen Mu, Yangze Li, Mu, Bingshen +13 · 7 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #Speech and dialogue systems #electronic engineering #information engineering
  29. FreeV: Free Lunch For Vocoders Through Pseudo Inversed Mel Filter
    2024/06/12 by Yuanjun Lv, Hai Li, Lv, Yuanjun +9 · 6 citations
    Engineering · #Artificial Immune Systems Applications #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Intravenous Infusion Technology and Safety #Nanomaterials and Printing Technologies #Sound (cs.SD) #electronic engineering #information engineering
  30. Controllable Emotion Transfer For End-to-End Speech Synthesis
    2020/11/17 by Tao Li, Shan Yang, Li, Tao +5 · 3 citations
    Computer Science · #Speech Recognition and Synthesis #Speech and Audio Processing #Advanced Data Compression Techniques
  31. Distinctive and Natural Speaker Anonymization via Singular Value Transformation-assisted Matrix
    2024/05/17 by Jixun Yao, Yao, Jixun, Qing Wang +7 · 6 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Electrical engineering #Natural Language Processing Techniques #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  32. DCCRN+: Channel-wise Subband DCCRN with SNR Estimation for Speech Enhancement
    2021/06/16 by Shubo Lv, Lv, Shubo, Yanxin Hu +5 · 3 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  33. E-chat: Emotion-sensitive Spoken Dialogue System with Large Language Models
    2023/12/31 by Hongfei Xue, Yuhao Liang, Xue, Hongfei +10 · 5 citations
    Computer Science · Psychology · #Audio and Speech Processing (eess.AS) #Emotion and Mood Recognition #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and dialogue systems #electronic engineering #information engineering
  34. PromptVC: Flexible Stylistic Voice Conversion in Latent Space Driven by Natural Language Prompts
    2023/09/17 by Jixun Yao, Yuguang Yang, Yao, Jixun +17 · 4 citations
    Computer Science · Medicine · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #Speech Recognition and Synthesis #Voice and Speech Disorders #electronic engineering #information engineering
  35. OSUM: Advancing Open Speech Understanding Models with Limited Resources in Academia
    2025/01/23 by Geng, Xuelong, Wei, Kun, Qijie Shao +34 · 11 citations
    Computer Science · #Natural Language Processing Techniques
  36. A Comparative Study on Speaker-attributed Automatic Speech Recognition in Multi-party Meetings
    2022/03/31 by Fan Yu, Yu, Fan, Zhihao Du +7 · 3 citations
    Computer Science · #Speech Recognition and Synthesis #Speech and dialogue systems #Natural Language Processing Techniques
  37. Attention-based End-to-End Models for Small-Footprint Keyword Spotting
    2018/03/29 by Changhao Shan, Junbo Zhang, Shan, Changhao +5 · 4 citations
    Computer Science · #Speech Recognition and Synthesis #Topic Modeling #Natural Language Processing Techniques
  38. StableVC: Style Controllable Zero-Shot Voice Conversion with Conditional Flow Matching
    2024/12/06 by Jixun Yao, Yao, Jixun, Yuguang Yan +11 · 7 citations
    Computer Science · #Speech Recognition and Synthesis #Speech and Audio Processing #Speech and dialogue systems
  39. DiAD: A Diffusion-based Framework for Multi-class Anomaly Detection
    2023/12/11 by Haoyang He, Jiangning Zhang, He, Haoyang +15 · 4 citations
    Biochemistry, Genetics and Molecular Biology · Computer Science · #Anomaly Detection Techniques and Applications #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Yersinia bacterium, plague, ectoparasites research
  40. Boundary and Context Aware Training for CIF-based Non-Autoregressive End-to-end ASR
    2021/04/10 by Fan Yu, Yu, Fan, Haoneng Luo +15 · 3 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  41. The ASRU 2019 Mandarin-English Code-Switching Speech Recognition Challenge: Open Datasets, Tracks, Methods and Results
    2020/07/12 by Xian Shi, Qiangze Feng, Shi, Xian +3 · 2 citations
    Computer Science · Psychology · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Phonetics and Phonology Research #Sound (cs.SD) #Speech Recognition and Synthesis #electronic engineering #information engineering
  42. Unveiling the Potential of LLM-Based ASR on Chinese Open-Source Datasets
    2024/05/03 by Xuelong Geng, Tianyi Xu, Geng, Xuelong +21 · 6 citations
    Computer Science · #Natural Language Processing Techniques
  43. Fine-grained Emotion Strength Transfer, Control and Prediction for Emotional Speech Synthesis
    2020/11/17 by Yi Lei, Lei, Yi, Shan Yang +3 · 2 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  44. SongEval: A Benchmark Dataset for Song Aesthetics Evaluation
    2025/05/16 by Jixun Yao, Yao, Jixun, Guobin Ma +20 · 9 citations
    Computer Science · #Artificial Intelligence in Games #Audio and Speech Processing (eess.AS) #FOS: Electrical engineering #Music Technology and Sound Studies #Music and Audio Processing #electronic engineering #information engineering
  45. AnyoneNet: Synchronized Speech and Talking Head Generation for Arbitrary Person
    2021/08/09 by Xinsheng Wang, Qicong Xie, Wang, Xinsheng +7 · 2 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #FOS: Electrical engineering #Face recognition and analysis #Generative Adversarial Networks and Image Synthesis #Sound (cs.SD) #Speech and Audio Processing #electronic engineering #information engineering
  46. Sensing Mutual Information with Random Signals in Gaussian Channels
    2023/11/13 by Lei Xie, Xie, Lei, Fan Liu +7 · 4 citations
    Computer Science · Engineering · #Distributed Sensor Networks and Detection Algorithms #FOS: Computer and information sciences #FOS: Electrical engineering #Indoor and Outdoor Localization Technologies #Information Theory (cs.IT) #Signal Processing (eess.SP) #Sparse and Compressive Sensing Techniques #electronic engineering #information engineering
  47. FlowSE: Efficient and High-Quality Speech Enhancement via Flow Matching
    2025/05/26 by Ziqian Wang, Wang, Ziqian, Xinfa Zhu +14 · 11 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Electrical engineering #Music and Audio Processing #Signal Processing (eess.SP) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  48. DiCLET-TTS: Diffusion Model based Cross-lingual Emotion Transfer for Text-to-Speech -- A Study between English and Mandarin
    2023/09/02 by Tao Li, Li, Tao, Chenxu Hu +13 · 3 citations
    Psychology · Computer Science · #Phonetics and Phonology Research #Speech Recognition and Synthesis #Sentiment Analysis and Opinion Mining
  49. Enriching Source Style Transfer in Recognition-Synthesis based Non-Parallel Voice Conversion
    2021/06/16 by Zhichao Wang, Xinyong Zhou, Wang, Zhichao +15 · 2 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  50. The second multi-channel multi-party meeting transcription challenge (M2MeT) 2.0): A benchmark for speaker-attributed ASR
    2023/09/24 by Yuhao Liang, Mohan Shi, Liang, Yuhao +24 · 3 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Topic Modeling #electronic engineering #information engineering
  51. VITS-Based Singing Voice Conversion Leveraging Whisper and multi-scale F0 Modeling
    2023/10/04 by Ziqian Ning, Yuepeng Jiang, Ning, Ziqian +7 · 3 citations
    Computer Science · #Speech Recognition and Synthesis #Speech and Audio Processing #Music and Audio Processing
  52. KALL-E:Autoregressive Speech Synthesis with Next-Distribution Prediction
    2024/12/22 by Xia, Kangxiang, Xinfa Zhu, Zhu, Xinfa +6 · 5 citations
    Computer Science · Psychology · #Speech Recognition and Synthesis #Speech and Audio Processing #Phonetics and Phonology Research
  53. UniSyn: An End-to-End Unified Model for Text-to-Speech and Singing Voice Synthesis
    2022/12/03 by Yi Lei, Shan Yang, Lei, Yi +11 · 2 citations
    Computer Science · #Speech Recognition and Synthesis #Music and Audio Processing #Speech and Audio Processing
  54. The NPU-Elevoc Personalized Speech Enhancement System for ICASSP2023 DNS Challenge
    2023/03/13 by Xiaopeng Yan, Yindi Yang, Yan, Xiaopeng +7 · 2 citations
    Computer Science · #Speech and Audio Processing #Speech Recognition and Synthesis #Music and Audio Processing
  55. Streaming Decoder-Only Automatic Speech Recognition with Discrete Speech Units: A Pilot Study
    2024/06/27 by Peikun Chen, Chen, Peikun, Sining Sun +7 · 3 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #Speech Recognition and Synthesis #electronic engineering #information engineering
  56. Security Vulnerabilities in Ethereum Smart Contracts: A Systematic Analysis
    2025/04/08 by Jasmine Wu, Lei Xie, Wu, Jixuan +3 · 6 citations
    Computer Science · #Blockchain Technology Applications and Security #Advanced Authentication Protocols Security #Digital Rights Management and Security
  57. Vec-Tok-VC+: Residual-enhanced Robust Zero-shot Voice Conversion with Progressive Constraints in a Dual-mode Training Strategy
    2024/06/14 by Linhan Ma, Ma, Linhan, Xinfa Zhu +13 · 3 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #Speech and dialogue systems #electronic engineering #information engineering
  58. MobileMamba: Lightweight Multi-Receptive Visual Mamba Network
    2024/11/24 by Haoyang He, Jiangning Zhang, He, Haoyang +17 · 3 citations
    Computer Science · #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Video Surveillance and Tracking Methods
  59. Inaudible Adversarial Perturbations for Targeted Attack in Speaker Recognition
    2020/05/21 by Qing Wang, Pengcheng Guo, Wang, Qing +3 · 1 citation
    Computer Science · #Adversarial Robustness in Machine Learning #Anomaly Detection Techniques and Applications #Audio and Speech Processing (eess.AS) #Digital Media Forensic Detection #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
  60. Optimizing voice conversion network with cycle consistency loss of\n speaker identity
    2020/11/17 by Hongqiang Du, Xiaohai Tian, Du, Hongqiang +5 · 1 citation
    Computer Science · #Speech Recognition and Synthesis #Speech and Audio Processing #Advanced Data Compression Techniques
  61. SQ-Whisper: Speaker-Querying based Whisper Model for Target-Speaker ASR
    2024/12/07 by Pengcheng Guo, Guo, Pengcheng, Xuankai Chang +7 · 3 citations
    Computer Science · #Speech Recognition and Synthesis #Speech and dialogue systems
  62. Statistical Parametric Speech Synthesis Using Generative Adversarial Networks Under A Multi-task Learning Framework
    2017/07/06 by Shan Yang, Lei Xie, Yang, Shan +11 · 1 citation
    Computer Science · #FOS: Computer and information sciences #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing
  63. StreamFlow: Streaming Flow Matching with Block-wise Guided Attention Mask for Speech Token Decoding
    2025/06/30 by Dake Guo, Jixun Yao, Guo, Dake +7 · 1 voice
    Computer Science · Engineering · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #cs.SD #eess.AS #electronic engineering #information engineering
  64. MsEmoTTS: Multi-scale emotion transfer, prediction, and control for emotional speech synthesis
    2022/01/17 by Yi Lei, Lei, Yi, Shan Yang +5 · 1 citation
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #Speech Recognition and Synthesis #electronic engineering #information engineering
  65. Zero-Shot Emotion Transfer For Cross-Lingual Speech Synthesis
    2023/10/06 by Yuke Li, Xinfa Zhu, Li, Yuke +11 · 2 citations
    Computer Science · #Speech Recognition and Synthesis #Speech and Audio Processing
  66. CaTT-KWS: A Multi-stage Customized Keyword Spotting Framework based on Cascaded Transducer-Transformer
    2022/07/04 by Zhanheng Yang, Sining Sun, Yang, Zhanheng +11 · 1 citation
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  67. WeKws: A production first small-footprint end-to-end Keyword Spotting Toolkit
    2022/10/30 by Jie Wang, Wang, Jie, Menglong Xu +11 · 1 citation
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #ICT in Developing Communities #Sound (cs.SD) #Speech and dialogue systems #Topic Modeling #electronic engineering #information engineering
  68. Multi-Speaker Expressive Speech Synthesis via Multiple Factors Decoupling
    2022/11/19 by Xinfa Zhu, Yi Lei, Zhu, Xinfa +9 · 1 citation
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  69. Improving Multimodal Emotion Recognition by Leveraging Acoustic Adaptation and Visual Alignment
    2024/09/08 by Zhixian Zhao, Zhao, Zhixian, Haifeng Chen +7 · 2 citations
    Psychology · #Audio and Speech Processing (eess.AS) #Emotion and Mood Recognition #FOS: Computer and information sciences #FOS: Electrical engineering #Human-Computer Interaction (cs.HC) #Sound (cs.SD) #electronic engineering #information engineering
  70. Approximate Message Passing-Enhanced Graph Neural Network for OTFS Data Detection
    2024/02/15 by Yuyi Mao, Zhuang, Wenhao, Mao, Yuyi +10 · 1 citation
    Engineering · #FOS: Computer and information sciences #FOS: Electrical engineering #Information Theory (cs.IT) #Optical Systems and Laser Technology #Signal Processing (eess.SP) #electronic engineering #information engineering
  71. StyleS2ST: Zero-shot Style Transfer for Direct Speech-to-speech Translation
    2023/05/28 by Kun Song, Yi Ren, Song, Kun +13 · 1 citation
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Topic Modeling #electronic engineering #information engineering
  72. MSM-VC: High-fidelity Source Style Transfer for Non-Parallel Voice Conversion by Multi-scale Style Modeling
    2023/09/03 by Zhichao Wang, Xinsheng Wang, Wang, Zhichao +11 · 1 citation
    Computer Science · Medicine · #Speech Recognition and Synthesis #Speech and Audio Processing #Voice and Speech Disorders
  73. ContextASR-Bench: A Massive Contextual Speech Recognition Benchmark
    2025/07/08 by He Wang, Wang, He, Linhan Ma +11 · 4 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Topic Modeling #electronic engineering #information engineering
  74. HiGNN-TTS: Hierarchical Prosody Modeling with Graph Neural Networks for Expressive Long-form TTS
    2023/09/25 by Dake Guo, Xinfa Zhu, Guo, Dake +11 · 1 citation
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Topic Modeling #electronic engineering #information engineering
  75. Linguistic-Acoustic Similarity Based Accent Shift for Accent Recognition
    2022/04/07 by Qijie Shao, Jinghao Yan, Shao, Qijie +11 · 1 citation
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  76. DiffRhythm+: Controllable and Flexible Full-Length Song Generation with Preference Optimization
    2025/07/17 by Chen, Huakang, Yu-rou JIANG, Jiang, Yuepeng +15 · 6 citations
    Computer Science · Engineering · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Human Motion and Animation #Music Technology and Sound Studies #Music and Audio Processing #Sound (cs.SD) #electronic engineering #information engineering
  77. SALT: Distinguishable Speaker Anonymization Through Latent Space Transformation
    2023/10/08 by Yuanjun Lv, Lv, Yuanjun, Jixun Yao +9 · 1 citation
    Computer Science · Medicine · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #Voice and Speech Disorders #electronic engineering #information engineering
  78. Cross-speaker Emotion Transfer Based On Prosody Compensation for End-to-End Speech Synthesis
    2022/07/04 by Tao Li, Xinsheng Wang, Li, Tao +9 · 1 citation
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  79. DualSep: A Light-weight dual-encoder convolutional recurrent network for real-time in-car speech separation
    2024/09/13 by Ziqian Wang, Wang, Ziqian, Jiayao Sun +9 · 2 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  80. Accent-VITS:accent transfer for end-to-end TTS
    2023/12/28 by Linhan Ma, Ma, Linhan, Yongmao Zhang +11 · 1 citation
    Computer Science · Psychology · Medicine · #Speech Recognition and Synthesis #Phonetics and Phonology Research #Voice and Speech Disorders
  81. Fine-grained Preference Optimization Improves Zero-shot Text-to-Speech
    2025/02/05 by Jixun Yao, Yuguang Yang, Yao, Jixun +13 · 2 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #Speech and dialogue systems #electronic engineering #information engineering
  82. Leveraging Acoustic Contextual Representation by Audio-textual Cross-modal Learning for Conversational ASR
    2022/07/03 by Kun Wei, Wei, Kun, Yike Zhang +7 · 1 citation
    Computer Science · Psychology · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Phonetics and Phonology Research #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  83. RaD-Net 2: A causal two-stage repairing and denoising speech enhancement network with knowledge distillation and complex axial self-attention
    2024/06/11 by Mingshuai Liu, Zhuangqi Chen, Liu, Mingshuai +13 · 1 citation
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  84. SongFormer: Scaling Music Structure Analysis with Heterogeneous Supervision
    2025/10/03 by C.X. Hao, Hao, Chunbo, Ruibin Yuan +10 · 2 citations
    Computer Science · #Music and Audio Processing #Speech Recognition and Synthesis #Speech and Audio Processing
  85. Towards Rehearsal-Free Multilingual ASR: A LoRA-based Case Study on Whisper
    2024/08/20 by Tianyi Xu, Xu, Tianyi, Kaixun Huang +11 · 1 citation
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Hate Speech and Cyberbullying Detection #Sound (cs.SD) #electronic engineering #information engineering
  86. Rapid and Safe Trajectory Planning over Diverse Scenes through Diffusion Composition
    2025/07/06 by Wule Mao, Zhouheng Li, Mao, Wule +7 · 2 citations
    Computer Science · Engineering · #Autonomous Vehicle Technology and Safety #Human Motion and Animation #Robotic Path Planning Algorithms #cs.RO #cs.SY #eess.SY
  87. A Transcription Prompt-based Efficient Audio Large Language Model for Robust Speech Recognition
    2024/08/18 by Yangze Li, Xiong Wang, Li, Yangze +9 · 1 citation
    Computer Science · #Speech Recognition and Synthesis #Speech and Audio Processing #Music and Audio Processing
  88. Leveraging LLM and Self-Supervised Training Models for Speech Recognition in Chinese Dialects: A Comparative Analysis
    2025/05/27 by Tianyi Xu, Xu, Tianyi, Hongjie Chen +14 · 2 citations
    Computer Science · #Speech Recognition and Synthesis
  89. Takin-VC: Expressive Zero-Shot Voice Conversion via Adaptive Hybrid Content Encoding and Enhanced Timbre Modeling
    2024/10/02 by Yuguang Yang, Yu Pan, Yang, Yuguang +15 · 1 citation
    Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #electronic engineering #information engineering
  90. YingMusic-Singer: Zero-shot Singing Voice Synthesis and Editing with Annotation-free Melody Guidance
    2025/12/04 by Junjie Zheng, Zheng, Junjie, Guobin Ma +11 · 1 citation
    Computer Science · #Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Music Technology and Sound Studies #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis
  91. WenetSpeech-Yue: A Large-scale Cantonese Speech Corpus with Multi-dimensional Annotation
    2025/09/04 by Longhao Li, Zhao Guo, Li, Longhao +33 · 3 citations
    Computer Science · #FOS: Computer and information sciences #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Topic Modeling
  92. Drop the beat! Freestyler for Accompaniment Conditioned Rapping Voice Generation
    2024/08/28 by Ziqian Ning, Ning, Ziqian, Shuai Wang +13 · 1 citation
    Computer Science · #Music Technology and Sound Studies #Speech and Audio Processing
  93. Sign Language Question Answering: A New Task, Benchmark, and Baseline for Sign Language Understanding
    2026/07/30 by Shiwei Gan, Lichen Wang, Xiao Liu +4
    Computer Science · #cs.AI #cs.CV
  94. Qwen-Music Technical Report
    2026/07/27 by Jin Xu, Kangdi Wang, Ruibin Yuan +24
    #cs.SD
  95. When Derived Measurements Mislead: Quantifying and Mitigating LLM Over-Trust with Privileged-Modality Reliability Evidence
    2026/07/30 by Zongheng Guo, Tao Chen, Tianli Li +6
    Computer Science · #cs.AI
  96. SGTP: Sampling-based Game-Theoretic Planning for Real-Time Multi-Vehicle Autonomous Racing
    2026/07/28 by Zhouheng Li, Fangguo Zhao, Mattia Piccinini +6
    #cs.RO
  97. SLT 2026 REAL-TSE Challenge: Real-world Target Speaker Extraction from Conversational Recordings
    2026/07/16 by Shuai Wang, Zihan Qian, Ke Zhang +9
    #eess.AS #cs.SD