Lei Xie
- Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens
2025/03/03 by Xinsheng Wang, Wang, Xinsheng, Mingqi Jiang +49 · 2 voices · 59 citations
Computer Science · #Natural Language Processing Techniques #Speech Recognition and Synthesis #Speech and dialogue systems #cs.AI #cs.SD #eess.AS
- DCCRN: Deep Complex Convolution Recurrent Network for Phase-Aware Speech Enhancement
2020/08/01 by Yanxin Hu, Yun Liu, Hu, Yanxin +15 · 34 citations
Computer Science · Neuroscience · #Speech and Audio Processing #Hearing Loss and Rehabilitation #Speech Recognition and Synthesis
- WenetSpeech: A 10000+ Hours Multi-domain Mandarin Corpus for Speech Recognition
2021/10/07 by Binbin Zhang, Zhang, Binbin, Hang Lv +21 · 30 citations
Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing
- Freeze-Omni: A Smart and Low Latency Speech-to-speech Dialogue Model with Frozen LLM
2024/11/01 by Xiong Wang, Wang, Xiong, Yangze Li +13 · 47 citations
Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and dialogue systems #Topic Modeling #electronic engineering #information engineering
- M2MeT: The ICASSP 2022 Multi-Channel Multi-Party Meeting Transcription Challenge
2021/10/14 by Fan Yu, Shiliang Zhang, Yu, Fan +21 · 17 citations
Computer Science · #Speech Recognition and Synthesis #Music and Audio Processing #Topic Modeling
- AISHELL-4: An Open Source Dataset for Speech Enhancement, Separation,\n Recognition and Speaker Diarization in Conference Scenario
2021/04/08 by Yihui Fu, Luyao Cheng, Fu, Yihui +22 · 16 citations
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- MambaAD: Exploring State Space Models for Multi-class Unsupervised Anomaly Detection
2024/04/09 by Haoyang He, Yuhu Bai, He, Haoyang +17 · 23 citations
Computer Science · Engineering · #Anomaly Detection Techniques and Applications #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Network Security and Intrusion Detection #Smart Grid Security and Resilience
- WeNet: Production oriented Streaming and Non-streaming End-to-End Speech Recognition Toolkit
2021/02/02 by Zhuoyuan Yao, Yao, Zhuoyuan, Di Wu +17 · 12 citations
Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #electronic engineering #information engineering
- Open Source MagicData-RAMC: A Rich Annotated Mandarin Conversational(RAMC) Speech Dataset
2022/03/31 by Zehui Yang, Yifan Chen, Yang, Zehui +21 · 11 citations
Computer Science · #Speech Recognition and Synthesis #Speech and dialogue systems #Topic Modeling
- SELM: Speech Enhancement Using Discrete Tokens and Language Models
2023/12/15 by Ziqian Wang, Xinfa Zhu, Wang, Ziqian +11 · 16 citations
Computer Science · Medicine · #Speech and Audio Processing #Speech Recognition and Synthesis #Voice and Speech Disorders
- Time Domain Audio Visual Speech Separation
2019/04/07 by Jian Wu, Yong Xu, Wu, Jian +11 · 7 citations
Computer Science · Engineering · #Advanced Adaptive Filtering Techniques #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech and Audio Processing #electronic engineering #information engineering
- WeNet 2.0: More Productive End-to-End Speech Recognition Toolkit
2022/03/29 by Binbin Zhang, Di Wu, Zhang, Binbin +17 · 9 citations
Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Topic Modeling #electronic engineering #information engineering
- Multi-band MelGAN: Faster Waveform Generation for High-Quality Text-to-Speech
2020/05/11 by Geng Yang, Shan Yang, Yang, Geng +9 · 7 citations
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- WenetSpeech4TTS: A 12,800-hour Mandarin TTS Corpus for Large Speech Generation Model Benchmark
2024/06/09 by Linhan Ma, Ma, Linhan, Dake Guo +16 · 15 citations
Computer Science · Psychology · #Audio and Speech Processing (eess.AS) #FOS: Electrical engineering #Phonetics and Phonology Research #Speech Recognition and Synthesis #Speech and dialogue systems #electronic engineering #information engineering
- Single-Codec: Single-Codebook Speech Codec towards High-Performance Speech Generation
2024/06/11 by Hanzhao Li, Liumeng Xue, Li, Hanzhao +15 · 14 citations
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Electrical engineering #Speech Recognition and Synthesis #Speech and dialogue systems #electronic engineering #information engineering
- Collaborative Sensing in Perceptive Mobile Networks: Opportunities and Challenges
2022/05/31 by Lei Xie, S. H. Song, Xie, Lei +5 · 7 citations
Computer Science · Engineering · #Energy Efficient Wireless Sensor Networks #Energy Harvesting in Wireless Networks #FOS: Computer and information sciences #FOS: Electrical engineering #Indoor and Outdoor Localization Technologies #Information Theory (cs.IT) #Signal Processing (eess.SP) #electronic engineering #information engineering
- GenSE: Generative Speech Enhancement via Language Models using Hierarchical Modeling
2025/02/05 by Jixun Yao, Hexin Liu, Yao, Jixun +9 · 14 citations
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #Speech and dialogue systems #electronic engineering #information engineering
- TFGAN: Time and Frequency Domain Based Generative Adversarial Network for High-fidelity Speech Synthesis
2020/11/24 by Qiao Tian, Yi Chen, Tian, Qiao +11 · 4 citations
Computer Science · #Speech Recognition and Synthesis #Speech and Audio Processing #Music and Audio Processing
- A Comprehensive Library for Benchmarking Multi-class Visual Anomaly Detection
2024/06/05 by Jiangning Zhang, Haoyang He, Zhang, Jiangning +17 · 8 citations
Computer Science · Medicine · #Anomaly Detection Techniques and Applications #COVID-19 diagnosis using AI #Computer Vision and Pattern Recognition (cs.CV) #Currency Recognition and Detection #FOS: Computer and information sciences
- METTS: Multilingual Emotional Text-to-Speech by Cross-speaker and Cross-lingual Emotion Transfer
2023/07/29 by Xinfa Zhu, Yi Lei, Zhu, Xinfa +11 · 6 citations
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Electrical engineering #Speech Recognition and Synthesis #electronic engineering #information engineering
- DSPGAN: a GAN-based universal vocoder for high-fidelity TTS by time-frequency domain supervision from DSP
2022/11/02 by Kun Song, Song, Kun, Yongmao Zhang +13 · 5 citations
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- PromptStyle: Controllable Style Transfer for Text-to-Speech with Natural Language Descriptions
2023/05/31 by Liu, Guanghou, Yongmao Zhang, Zhang, Yongmao +10 · 5 citations
Computer Science · #Speech Recognition and Synthesis #Speech and Audio Processing #Music and Audio Processing
- Molecular Mechanics-Driven Graph Neural Network with Multiplex Graph for Molecular Structures
2020/11/15 by Shuo Zhang, Yang Liu, Zhang, Shuo +3 · 4 citations
Materials Science · Computer Science · Biochemistry, Genetics and Molecular Biology · #Machine Learning in Materials Science #Computational Drug Discovery Methods #Protein Structure and Dynamics
- Summary On The ICASSP 2022 Multi-Channel Multi-Party Meeting Transcription Grand Challenge
2022/02/08 by Fan Yu, Yu, Fan, Shiliang Zhang +29 · 4 citations
Computer Science · #Speech Recognition and Synthesis #Music and Audio Processing #Speech and Audio Processing
- Conversational End-to-End TTS for Voice Agent
2020/05/21 by Haohan Guo, Guo, Haohan, Shaofei Zhang +7 · 3 citations
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Topic Modeling #electronic engineering #information engineering
- Perceptive Mobile Network with Distributed Target Monitoring Terminals: Leaking Communication Energy for Sensing
2021/12/29 by Lei Xie, Xie, Lei, Peilan Wang +5 · 4 citations
Engineering · #Energy Harvesting in Wireless Networks #FOS: Computer and information sciences #FOS: Electrical engineering #Full-Duplex Wireless Communications #Indoor and Outdoor Localization Technologies #Information Theory (cs.IT) #Signal Processing (eess.SP) #electronic engineering #information engineering
- Uformer: A Unet based dilated complex & real dual-path conformer network for simultaneous speech enhancement and dereverberation
2021/11/11 by Yihui Fu, Fu, Yihui, Yun Liu +11 · 4 citations
Computer Science · Health Professions · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Infant Health and Development #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- MMGER: Multi-modal and Multi-granularity Generative Error Correction with LLM for Joint Accent and Speech Recognition
2024/05/06 by Bingshen Mu, Yangze Li, Mu, Bingshen +13 · 7 citations
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #Speech and dialogue systems #electronic engineering #information engineering
- FreeV: Free Lunch For Vocoders Through Pseudo Inversed Mel Filter
2024/06/12 by Yuanjun Lv, Hai Li, Lv, Yuanjun +9 · 6 citations
Engineering · #Artificial Immune Systems Applications #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Intravenous Infusion Technology and Safety #Nanomaterials and Printing Technologies #Sound (cs.SD) #electronic engineering #information engineering
- Controllable Emotion Transfer For End-to-End Speech Synthesis
2020/11/17 by Tao Li, Shan Yang, Li, Tao +5 · 3 citations
Computer Science · #Speech Recognition and Synthesis #Speech and Audio Processing #Advanced Data Compression Techniques
- Distinctive and Natural Speaker Anonymization via Singular Value Transformation-assisted Matrix
2024/05/17 by Jixun Yao, Yao, Jixun, Qing Wang +7 · 6 citations
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Electrical engineering #Natural Language Processing Techniques #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- DCCRN+: Channel-wise Subband DCCRN with SNR Estimation for Speech Enhancement
2021/06/16 by Shubo Lv, Lv, Shubo, Yanxin Hu +5 · 3 citations
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- E-chat: Emotion-sensitive Spoken Dialogue System with Large Language Models
2023/12/31 by Hongfei Xue, Yuhao Liang, Xue, Hongfei +10 · 5 citations
Computer Science · Psychology · #Audio and Speech Processing (eess.AS) #Emotion and Mood Recognition #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and dialogue systems #electronic engineering #information engineering
- PromptVC: Flexible Stylistic Voice Conversion in Latent Space Driven by Natural Language Prompts
2023/09/17 by Jixun Yao, Yuguang Yang, Yao, Jixun +17 · 4 citations
Computer Science · Medicine · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #Speech Recognition and Synthesis #Voice and Speech Disorders #electronic engineering #information engineering
- OSUM: Advancing Open Speech Understanding Models with Limited Resources in Academia
2025/01/23 by Geng, Xuelong, Wei, Kun, Qijie Shao +34 · 11 citations
Computer Science · #Natural Language Processing Techniques
- A Comparative Study on Speaker-attributed Automatic Speech Recognition in Multi-party Meetings
2022/03/31 by Fan Yu, Yu, Fan, Zhihao Du +7 · 3 citations
Computer Science · #Speech Recognition and Synthesis #Speech and dialogue systems #Natural Language Processing Techniques
- Attention-based End-to-End Models for Small-Footprint Keyword Spotting
2018/03/29 by Changhao Shan, Junbo Zhang, Shan, Changhao +5 · 4 citations
Computer Science · #Speech Recognition and Synthesis #Topic Modeling #Natural Language Processing Techniques
- StableVC: Style Controllable Zero-Shot Voice Conversion with Conditional Flow Matching
2024/12/06 by Jixun Yao, Yao, Jixun, Yuguang Yan +11 · 7 citations
Computer Science · #Speech Recognition and Synthesis #Speech and Audio Processing #Speech and dialogue systems
- DiAD: A Diffusion-based Framework for Multi-class Anomaly Detection
2023/12/11 by Haoyang He, Jiangning Zhang, He, Haoyang +15 · 4 citations
Biochemistry, Genetics and Molecular Biology · Computer Science · #Anomaly Detection Techniques and Applications #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Yersinia bacterium, plague, ectoparasites research
- Boundary and Context Aware Training for CIF-based Non-Autoregressive End-to-end ASR
2021/04/10 by Fan Yu, Yu, Fan, Haoneng Luo +15 · 3 citations
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- The ASRU 2019 Mandarin-English Code-Switching Speech Recognition Challenge: Open Datasets, Tracks, Methods and Results
2020/07/12 by Xian Shi, Qiangze Feng, Shi, Xian +3 · 2 citations
Computer Science · Psychology · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Phonetics and Phonology Research #Sound (cs.SD) #Speech Recognition and Synthesis #electronic engineering #information engineering
- Unveiling the Potential of LLM-Based ASR on Chinese Open-Source Datasets
2024/05/03 by Xuelong Geng, Tianyi Xu, Geng, Xuelong +21 · 6 citations
Computer Science · #Natural Language Processing Techniques
- Fine-grained Emotion Strength Transfer, Control and Prediction for Emotional Speech Synthesis
2020/11/17 by Yi Lei, Lei, Yi, Shan Yang +3 · 2 citations
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- SongEval: A Benchmark Dataset for Song Aesthetics Evaluation
2025/05/16 by Jixun Yao, Yao, Jixun, Guobin Ma +20 · 9 citations
Computer Science · #Artificial Intelligence in Games #Audio and Speech Processing (eess.AS) #FOS: Electrical engineering #Music Technology and Sound Studies #Music and Audio Processing #electronic engineering #information engineering
- AnyoneNet: Synchronized Speech and Talking Head Generation for Arbitrary Person
2021/08/09 by Xinsheng Wang, Qicong Xie, Wang, Xinsheng +7 · 2 citations
Computer Science · #Audio and Speech Processing (eess.AS) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #FOS: Electrical engineering #Face recognition and analysis #Generative Adversarial Networks and Image Synthesis #Sound (cs.SD) #Speech and Audio Processing #electronic engineering #information engineering
- Sensing Mutual Information with Random Signals in Gaussian Channels
2023/11/13 by Lei Xie, Xie, Lei, Fan Liu +7 · 4 citations
Computer Science · Engineering · #Distributed Sensor Networks and Detection Algorithms #FOS: Computer and information sciences #FOS: Electrical engineering #Indoor and Outdoor Localization Technologies #Information Theory (cs.IT) #Signal Processing (eess.SP) #Sparse and Compressive Sensing Techniques #electronic engineering #information engineering
- FlowSE: Efficient and High-Quality Speech Enhancement via Flow Matching
2025/05/26 by Ziqian Wang, Wang, Ziqian, Xinfa Zhu +14 · 11 citations
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Electrical engineering #Music and Audio Processing #Signal Processing (eess.SP) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- DiCLET-TTS: Diffusion Model based Cross-lingual Emotion Transfer for Text-to-Speech -- A Study between English and Mandarin
2023/09/02 by Tao Li, Li, Tao, Chenxu Hu +13 · 3 citations
Psychology · Computer Science · #Phonetics and Phonology Research #Speech Recognition and Synthesis #Sentiment Analysis and Opinion Mining
- Enriching Source Style Transfer in Recognition-Synthesis based Non-Parallel Voice Conversion
2021/06/16 by Zhichao Wang, Xinyong Zhou, Wang, Zhichao +15 · 2 citations
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- The second multi-channel multi-party meeting transcription challenge (M2MeT) 2.0): A benchmark for speaker-attributed ASR
2023/09/24 by Yuhao Liang, Mohan Shi, Liang, Yuhao +24 · 3 citations
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Topic Modeling #electronic engineering #information engineering
- VITS-Based Singing Voice Conversion Leveraging Whisper and multi-scale F0 Modeling
2023/10/04 by Ziqian Ning, Yuepeng Jiang, Ning, Ziqian +7 · 3 citations
Computer Science · #Speech Recognition and Synthesis #Speech and Audio Processing #Music and Audio Processing
- KALL-E:Autoregressive Speech Synthesis with Next-Distribution Prediction
2024/12/22 by Xia, Kangxiang, Xinfa Zhu, Zhu, Xinfa +6 · 5 citations
Computer Science · Psychology · #Speech Recognition and Synthesis #Speech and Audio Processing #Phonetics and Phonology Research
- UniSyn: An End-to-End Unified Model for Text-to-Speech and Singing Voice Synthesis
2022/12/03 by Yi Lei, Shan Yang, Lei, Yi +11 · 2 citations
Computer Science · #Speech Recognition and Synthesis #Music and Audio Processing #Speech and Audio Processing
- The NPU-Elevoc Personalized Speech Enhancement System for ICASSP2023 DNS Challenge
2023/03/13 by Xiaopeng Yan, Yindi Yang, Yan, Xiaopeng +7 · 2 citations
Computer Science · #Speech and Audio Processing #Speech Recognition and Synthesis #Music and Audio Processing
- Streaming Decoder-Only Automatic Speech Recognition with Discrete Speech Units: A Pilot Study
2024/06/27 by Peikun Chen, Chen, Peikun, Sining Sun +7 · 3 citations
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #Speech Recognition and Synthesis #electronic engineering #information engineering
- Security Vulnerabilities in Ethereum Smart Contracts: A Systematic Analysis
2025/04/08 by Jasmine Wu, Lei Xie, Wu, Jixuan +3 · 6 citations
Computer Science · #Blockchain Technology Applications and Security #Advanced Authentication Protocols Security #Digital Rights Management and Security
- Vec-Tok-VC+: Residual-enhanced Robust Zero-shot Voice Conversion with Progressive Constraints in a Dual-mode Training Strategy
2024/06/14 by Linhan Ma, Ma, Linhan, Xinfa Zhu +13 · 3 citations
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #Speech and dialogue systems #electronic engineering #information engineering
- MobileMamba: Lightweight Multi-Receptive Visual Mamba Network
2024/11/24 by Haoyang He, Jiangning Zhang, He, Haoyang +17 · 3 citations
Computer Science · #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Video Surveillance and Tracking Methods
- Inaudible Adversarial Perturbations for Targeted Attack in Speaker Recognition
2020/05/21 by Qing Wang, Pengcheng Guo, Wang, Qing +3 · 1 citation
Computer Science · #Adversarial Robustness in Machine Learning #Anomaly Detection Techniques and Applications #Audio and Speech Processing (eess.AS) #Digital Media Forensic Detection #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
- Optimizing voice conversion network with cycle consistency loss of\n speaker identity
2020/11/17 by Hongqiang Du, Xiaohai Tian, Du, Hongqiang +5 · 1 citation
Computer Science · #Speech Recognition and Synthesis #Speech and Audio Processing #Advanced Data Compression Techniques
- SQ-Whisper: Speaker-Querying based Whisper Model for Target-Speaker ASR
2024/12/07 by Pengcheng Guo, Guo, Pengcheng, Xuankai Chang +7 · 3 citations
Computer Science · #Speech Recognition and Synthesis #Speech and dialogue systems
- Statistical Parametric Speech Synthesis Using Generative Adversarial Networks Under A Multi-task Learning Framework
2017/07/06 by Shan Yang, Lei Xie, Yang, Shan +11 · 1 citation
Computer Science · #FOS: Computer and information sciences #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing
- StreamFlow: Streaming Flow Matching with Block-wise Guided Attention Mask for Speech Token Decoding
2025/06/30 by Dake Guo, Jixun Yao, Guo, Dake +7 · 1 voice
Computer Science · Engineering · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #cs.SD #eess.AS #electronic engineering #information engineering
- MsEmoTTS: Multi-scale emotion transfer, prediction, and control for emotional speech synthesis
2022/01/17 by Yi Lei, Lei, Yi, Shan Yang +5 · 1 citation
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #Speech Recognition and Synthesis #electronic engineering #information engineering
- Zero-Shot Emotion Transfer For Cross-Lingual Speech Synthesis
2023/10/06 by Yuke Li, Xinfa Zhu, Li, Yuke +11 · 2 citations
Computer Science · #Speech Recognition and Synthesis #Speech and Audio Processing
- CaTT-KWS: A Multi-stage Customized Keyword Spotting Framework based on Cascaded Transducer-Transformer
2022/07/04 by Zhanheng Yang, Sining Sun, Yang, Zhanheng +11 · 1 citation
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- WeKws: A production first small-footprint end-to-end Keyword Spotting Toolkit
2022/10/30 by Jie Wang, Wang, Jie, Menglong Xu +11 · 1 citation
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #ICT in Developing Communities #Sound (cs.SD) #Speech and dialogue systems #Topic Modeling #electronic engineering #information engineering
- Multi-Speaker Expressive Speech Synthesis via Multiple Factors Decoupling
2022/11/19 by Xinfa Zhu, Yi Lei, Zhu, Xinfa +9 · 1 citation
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- Improving Multimodal Emotion Recognition by Leveraging Acoustic Adaptation and Visual Alignment
2024/09/08 by Zhixian Zhao, Zhao, Zhixian, Haifeng Chen +7 · 2 citations
Psychology · #Audio and Speech Processing (eess.AS) #Emotion and Mood Recognition #FOS: Computer and information sciences #FOS: Electrical engineering #Human-Computer Interaction (cs.HC) #Sound (cs.SD) #electronic engineering #information engineering
- Approximate Message Passing-Enhanced Graph Neural Network for OTFS Data Detection
2024/02/15 by Yuyi Mao, Zhuang, Wenhao, Mao, Yuyi +10 · 1 citation
Engineering · #FOS: Computer and information sciences #FOS: Electrical engineering #Information Theory (cs.IT) #Optical Systems and Laser Technology #Signal Processing (eess.SP) #electronic engineering #information engineering
- StyleS2ST: Zero-shot Style Transfer for Direct Speech-to-speech Translation
2023/05/28 by Kun Song, Yi Ren, Song, Kun +13 · 1 citation
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Topic Modeling #electronic engineering #information engineering
- MSM-VC: High-fidelity Source Style Transfer for Non-Parallel Voice Conversion by Multi-scale Style Modeling
2023/09/03 by Zhichao Wang, Xinsheng Wang, Wang, Zhichao +11 · 1 citation
Computer Science · Medicine · #Speech Recognition and Synthesis #Speech and Audio Processing #Voice and Speech Disorders
- ContextASR-Bench: A Massive Contextual Speech Recognition Benchmark
2025/07/08 by He Wang, Wang, He, Linhan Ma +11 · 4 citations
Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Topic Modeling #electronic engineering #information engineering
- HiGNN-TTS: Hierarchical Prosody Modeling with Graph Neural Networks for Expressive Long-form TTS
2023/09/25 by Dake Guo, Xinfa Zhu, Guo, Dake +11 · 1 citation
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Topic Modeling #electronic engineering #information engineering
- Linguistic-Acoustic Similarity Based Accent Shift for Accent Recognition
2022/04/07 by Qijie Shao, Jinghao Yan, Shao, Qijie +11 · 1 citation
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- DiffRhythm+: Controllable and Flexible Full-Length Song Generation with Preference Optimization
2025/07/17 by Chen, Huakang, Yu-rou JIANG, Jiang, Yuepeng +15 · 6 citations
Computer Science · Engineering · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Human Motion and Animation #Music Technology and Sound Studies #Music and Audio Processing #Sound (cs.SD) #electronic engineering #information engineering
- SALT: Distinguishable Speaker Anonymization Through Latent Space Transformation
2023/10/08 by Yuanjun Lv, Lv, Yuanjun, Jixun Yao +9 · 1 citation
Computer Science · Medicine · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #Voice and Speech Disorders #electronic engineering #information engineering
- Cross-speaker Emotion Transfer Based On Prosody Compensation for End-to-End Speech Synthesis
2022/07/04 by Tao Li, Xinsheng Wang, Li, Tao +9 · 1 citation
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- DualSep: A Light-weight dual-encoder convolutional recurrent network for real-time in-car speech separation
2024/09/13 by Ziqian Wang, Wang, Ziqian, Jiayao Sun +9 · 2 citations
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- Accent-VITS:accent transfer for end-to-end TTS
2023/12/28 by Linhan Ma, Ma, Linhan, Yongmao Zhang +11 · 1 citation
Computer Science · Psychology · Medicine · #Speech Recognition and Synthesis #Phonetics and Phonology Research #Voice and Speech Disorders
- Fine-grained Preference Optimization Improves Zero-shot Text-to-Speech
2025/02/05 by Jixun Yao, Yuguang Yang, Yao, Jixun +13 · 2 citations
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #Speech and dialogue systems #electronic engineering #information engineering
- Leveraging Acoustic Contextual Representation by Audio-textual Cross-modal Learning for Conversational ASR
2022/07/03 by Kun Wei, Wei, Kun, Yike Zhang +7 · 1 citation
Computer Science · Psychology · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Phonetics and Phonology Research #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- RaD-Net 2: A causal two-stage repairing and denoising speech enhancement network with knowledge distillation and complex axial self-attention
2024/06/11 by Mingshuai Liu, Zhuangqi Chen, Liu, Mingshuai +13 · 1 citation
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- SongFormer: Scaling Music Structure Analysis with Heterogeneous Supervision
2025/10/03 by C.X. Hao, Hao, Chunbo, Ruibin Yuan +10 · 2 citations
Computer Science · #Music and Audio Processing #Speech Recognition and Synthesis #Speech and Audio Processing
- Towards Rehearsal-Free Multilingual ASR: A LoRA-based Case Study on Whisper
2024/08/20 by Tianyi Xu, Xu, Tianyi, Kaixun Huang +11 · 1 citation
Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Hate Speech and Cyberbullying Detection #Sound (cs.SD) #electronic engineering #information engineering
- Rapid and Safe Trajectory Planning over Diverse Scenes through Diffusion Composition
2025/07/06 by Wule Mao, Zhouheng Li, Mao, Wule +7 · 2 citations
Computer Science · Engineering · #Autonomous Vehicle Technology and Safety #Human Motion and Animation #Robotic Path Planning Algorithms #cs.RO #cs.SY #eess.SY
- A Transcription Prompt-based Efficient Audio Large Language Model for Robust Speech Recognition
2024/08/18 by Yangze Li, Xiong Wang, Li, Yangze +9 · 1 citation
Computer Science · #Speech Recognition and Synthesis #Speech and Audio Processing #Music and Audio Processing
- Leveraging LLM and Self-Supervised Training Models for Speech Recognition in Chinese Dialects: A Comparative Analysis
2025/05/27 by Tianyi Xu, Xu, Tianyi, Hongjie Chen +14 · 2 citations
Computer Science · #Speech Recognition and Synthesis
- Takin-VC: Expressive Zero-Shot Voice Conversion via Adaptive Hybrid Content Encoding and Enhanced Timbre Modeling
2024/10/02 by Yuguang Yang, Yu Pan, Yang, Yuguang +15 · 1 citation
Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #electronic engineering #information engineering
- YingMusic-Singer: Zero-shot Singing Voice Synthesis and Editing with Annotation-free Melody Guidance
2025/12/04 by Junjie Zheng, Zheng, Junjie, Guobin Ma +11 · 1 citation
Computer Science · #Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Music Technology and Sound Studies #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis
- WenetSpeech-Yue: A Large-scale Cantonese Speech Corpus with Multi-dimensional Annotation
2025/09/04 by Longhao Li, Zhao Guo, Li, Longhao +33 · 3 citations
Computer Science · #FOS: Computer and information sciences #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Topic Modeling
- Drop the beat! Freestyler for Accompaniment Conditioned Rapping Voice Generation
2024/08/28 by Ziqian Ning, Ning, Ziqian, Shuai Wang +13 · 1 citation
Computer Science · #Music Technology and Sound Studies #Speech and Audio Processing
- Sign Language Question Answering: A New Task, Benchmark, and Baseline for Sign Language Understanding
2026/07/30 by Shiwei Gan, Lichen Wang, Xiao Liu +4
Computer Science · #cs.AI #cs.CV
- Qwen-Music Technical Report
2026/07/27 by Jin Xu, Kangdi Wang, Ruibin Yuan +24
#cs.SD
- When Derived Measurements Mislead: Quantifying and Mitigating LLM Over-Trust with Privileged-Modality Reliability Evidence
2026/07/30 by Zongheng Guo, Tao Chen, Tianli Li +6
Computer Science · #cs.AI
- SGTP: Sampling-based Game-Theoretic Planning for Real-Time Multi-Vehicle Autonomous Racing
2026/07/28 by Zhouheng Li, Fangguo Zhao, Mattia Piccinini +6
#cs.RO
- SLT 2026 REAL-TSE Challenge: Real-world Target Speaker Extraction from Conversational Recordings
2026/07/16 by Shuai Wang, Zihan Qian, Ke Zhang +9
#eess.AS #cs.SD