Wen, Zhengqi
- ADD 2022: the First Audio Deep Synthesis Detection Challenge
2022/02/17 by Jiangyan Yi, Ruibo Fu, Yi, Jiangyan +36 · 23 citations
Computer Science · #Audio and Speech Processing (eess.AS) #Digital Media Forensic Detection #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Music and Audio Processing #Sound (cs.SD) #Speech and Audio Processing #electronic engineering #information engineering
- ADD 2023: the Second Audio Deepfake Detection Challenge
2023/05/23 by Jiangyan Yi, Yi, Jiangyan, Jianhua Tao +33 · 25 citations
Computer Science · #Audio and Speech Processing (eess.AS) #Digital Media Forensic Detection #FOS: Computer and information sciences #FOS: Electrical engineering #Generative Adversarial Networks and Image Synthesis #Music and Audio Processing #Sound (cs.SD) #electronic engineering #information engineering
- The Codecfake Dataset and Countermeasures for the Universally Detection of Deepfake Audio
2024/05/08 by Yuankun Xie, Xie, Yuankun, Yi Lu +21 · 12 citations
Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Digital Media Forensic Detection #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
- Mixture of Experts Fusion for Fake Audio Detection Using Frozen wav2vec 2.0
2024/09/18 by Wang, Zhiyong, Fu, Ruibo, Wen, Zhengqi +10 · 9 citations
#Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
- TraceableSpeech: Towards Proactively Traceable Text-to-Speech with Watermarking
2024/06/07 by Zhou, Junzuo, Yi, Jiangyan, Wang, Tao +5 · 7 citations
#Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
- Gated Recurrent Fusion with Joint Training Framework for Robust End-to-End Speech Recognition
2020/11/09 by Fan, Cunhang, Yi, Jiangyan, Tao, Jianhua +3 · 2 citations
#Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
- Fast End-to-End Speech Recognition via Non-Autoregressive Models and Cross-Modal Knowledge Transferring from BERT
2021/02/15 by Ye Bai, Bai, Ye, Jiangyan Yi +9 · 2 citations
Computer Science · #Speech Recognition and Synthesis #Music and Audio Processing #Natural Language Processing Techniques
- FSR: Accelerating the Inference Process of Transducer-Based Models by Applying Fast-Skip Regularization
2021/04/07 by Zhengkun Tian, Tian, Zhengkun, Jiangyan Yi +9 · 2 citations
Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- Spike-Triggered Non-Autoregressive Transformer for End-to-End Speech Recognition
2020/05/16 by Zhengkun Tian, Tian, Zhengkun, Jiangyan Yi +9 · 2 citations
Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #electronic engineering #information engineering
- TemplateRL: Structured Template-Guided Reinforcement Learning for LLM Reasoning
2025/05/21 by Jinyang Wu, Chan-Yu Liao, Wu, Jinyang +13 · 9 citations
Decision Sciences · #Complex Systems and Decision Making #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG)
- AStar: Boosting Multimodal Reasoning with Automated Structured Thinking
2025/02/04 by Jinyang Wu, Mingkuan Feng, Wu, Jinyang +14 · 4 citations
Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Natural Language Processing Techniques #Semantic Web and Ontologies #Speech and dialogue systems
- VQ-CTAP: Cross-Modal Fine-Grained Sequence Representation Learning for Speech Processing
2024/08/11 by Chunyu Qiang, Geng Wang, Qiang, Chunyu +26 · 5 citations
Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and dialogue systems #electronic engineering #information engineering
- Towards Fine-Grained Prosody Control for Voice Conversion
2019/10/24 by Zheng Lian, Zhengqi Wen, Lian, Zheng +1 · 1 citation
Computer Science · #Speech Recognition and Synthesis #Music and Audio Processing #Speech and Audio Processing
- Learn Spelling from Teachers: Transferring Knowledge from Language Models to Sequence-to-Sequence Speech Recognition
2019/07/13 by Ye Bai, Bai, Ye, Jiangyan Yi +7 · 1 citation
Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Natural Language Processing Techniques #Speech Recognition and Synthesis #electronic engineering #information engineering
- Integrating Knowledge into End-to-End Speech Recognition from External Text-Only Data
2019/12/04 by Ye Bai, Bai, Ye, Jiangyan Yi +9 · 1 citation
Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Topic Modeling #electronic engineering #information engineering
- Listen Attentively, and Spell Once: Whole Sentence Generation via a Non-Autoregressive Architecture for Low-Latency Speech Recognition
2020/05/11 by Ye Bai, Jiangyan Yi, Bai, Ye +9 · 1 citation
Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Speech Recognition and Synthesis #Topic Modeling #electronic engineering #information engineering
- Generalized Source Tracing: Detecting Novel Audio Deepfake Algorithm with Real Emphasis and Fake Dispersion Strategy
2024/06/05 by Yuankun Xie, Xie, Yuankun, Ruibo Fu +13 · 2 citations
Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Digital Media Forensic Detection #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech and Audio Processing #electronic engineering #information engineering
- Codecfake: An Initial Dataset for Detecting LLM-based Deepfake Audio
2024/06/12 by Lu, Yi, Xie, Yuankun, Fu, Ruibo +9 · 2 citations
#Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
- Singing-Tacotron: Global duration control attention and dynamic filter for End-to-end singing voice synthesis
2022/02/16 by Wang, Tao, Fu, Ruibo, Yi, Jiangyan +2 · 1 citation
#Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Sound (cs.SD) #electronic engineering #information engineering
- Exploring the Role of Audio in Multimodal Misinformation Detection
2024/08/22 by Yukun Liu, Liu, Moyang, Liu, Yukun +10 · 2 citations
Social Sciences · #Misinformation and Its Impacts
- MDPE: A Multimodal Deception Dataset with Personality and Emotional Characteristics
2024/07/17 by Cong Cai, Cai, Cong, Shan Liang +25 · 2 citations
Computer Science · Psychology · Social Sciences · #Cybercrime and Law Enforcement Studies #Deception detection and forensic psychology #Crime Patterns and Interventions
- Emotion Selectable End-to-End Text-based Speech Editing
2022/12/20 by Wang, Tao, Yi, Jiangyan, Fu, Ruibo +3 · 1 citation
#Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Sound (cs.SD) #electronic engineering #information engineering
- Learning From Yourself: A Self-Distillation Method for Fake Speech Detection
2023/03/02 by Cunhang Fan, Xue, Jun, Jiangyan Yi +10 · 1 citation
Computer Science · #Audio and Speech Processing (eess.AS) #Digital Media Forensic Detection #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Multimedia (cs.MM) #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- ALLM4ADD: Unlocking the Capabilities of Audio Large Language Models for Audio Deepfake Detection
2025/05/16 by Hao Gu, Jiangyan Yi, Gu, Hao +15 · 3 citations
Computer Science · #Generative Adversarial Networks and Image Synthesis #Speech Recognition and Synthesis #Music and Audio Processing
- Genuine-Focused Learning using Mask AutoEncoder for Generalized Fake Audio Detection
2024/06/05 by Wang, Xiaopeng, Fu, Ruibo, Wen, Zhengqi +9 · 2 citations
#Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
- DReSS: Data-driven Regularized Structured Streamlining for Large Language Models
2025/01/29 by Mingkuan Feng, Feng, Mingkuan, Jinyang Wu +13 · 1 citation
Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Natural Language Processing Techniques #Topic Modeling
- ASRRL-TTS: Agile Speaker Representation Reinforcement Learning for Text-to-Speech Speaker Adaptation
2024/07/07 by Ruibo Fu, Xin Qi, Fu, Ruibo +23 · 1 citation
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- Two-Stage Regularization-Based Structured Pruning for LLMs
2025/05/23 by Feng, Mingkuan, Wu, Jinyang, Liu, Siyuan +6 · 1 citation
#Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG)
- PPPR: Portable Plug-in Prompt Refiner for Text to Audio Generation
2024/06/07 by Shi, Shuchen, Fu, Ruibo, Wen, Zhengqi +10 · 1 citation
#Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
- Does Current Deepfake Audio Detection Model Effectively Detect ALM-based Deepfake Audio?
2024/08/20 by Yuankun Xie, Xie, Yuankun, Chenxu Xiong +21 · 1 citation
Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Digital Media Forensic Detection #FOS: Computer and information sciences #FOS: Electrical engineering #Image and Signal Denoising Methods #Music and Audio Processing #Sound (cs.SD) #electronic engineering #information engineering
- M3ANet: Multi-scale and Multi-Modal Alignment Network for Brain-Assisted Target Speaker Extraction
2025/05/31 by Cunhang Fan, Ying Chen, Fan, Cunhang +14 · 2 citations
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- Neural Codec Source Tracing: Toward Comprehensive Attribution in Open-Set Condition
2025/01/11 by Xie, Yuankun, Wang, Xiaopeng, Wang, Zhiyong +7 · 1 citation
#Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
- DPI-TTS: Directional Patch Interaction for Fast-Converging and Style Temporal Modeling in Text-to-Speech
2024/09/18 by Qi, Xin, Fu, Ruibo, Wen, Zhengqi +12 · 1 citation
#Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
- ImViD: Immersive Volumetric Videos for Enhanced VR Engagement
2025/03/18 by Yang, Zhengxian, Pan, Shi, Wang, Shengqi +7 · 1 citation
#Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences
- SecoustiCodec: Cross-Modal Aligned Streaming Single-Codecbook Speech Codec
2025/08/04 by Qiang, Chunyu, Wang, Haoyu, Gong, Cheng +10 · 3 citations
#Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
- RadialRouter: Structured Representation for Efficient and Robust Large Language Models Routing
2025/06/04 by Ruihan Jin, Pengpeng Shao, Jin, Ruihan +11 · 3 citations
Computer Science · Physics and Astronomy · #Natural Language Processing Techniques #Complex Network Analysis Techniques #Advanced Graph Neural Networks
- P2Mark: Plug-and-play Parameter-level Watermarking for Neural Speech Generation
2025/04/07 by Ren, Yong, Yi, Jiangyan, Wang, Tao +7 · 1 citation
#FOS: Computer and information sciences #Sound (cs.SD)