vix.ing · top · new · best · stats · spec

Wen, Zhengqi

  1. ADD 2022: the First Audio Deep Synthesis Detection Challenge
    2022/02/17 by Jiangyan Yi, Ruibo Fu, Yi, Jiangyan +36 · 23 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #Digital Media Forensic Detection #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Music and Audio Processing #Sound (cs.SD) #Speech and Audio Processing #electronic engineering #information engineering
  2. ADD 2023: the Second Audio Deepfake Detection Challenge
    2023/05/23 by Jiangyan Yi, Yi, Jiangyan, Jianhua Tao +33 · 25 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #Digital Media Forensic Detection #FOS: Computer and information sciences #FOS: Electrical engineering #Generative Adversarial Networks and Image Synthesis #Music and Audio Processing #Sound (cs.SD) #electronic engineering #information engineering
  3. The Codecfake Dataset and Countermeasures for the Universally Detection of Deepfake Audio
    2024/05/08 by Yuankun Xie, Xie, Yuankun, Yi Lu +21 · 12 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Digital Media Forensic Detection #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
  4. Mixture of Experts Fusion for Fake Audio Detection Using Frozen wav2vec 2.0
    2024/09/18 by Wang, Zhiyong, Fu, Ruibo, Wen, Zhengqi +10 · 9 citations
    #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
  5. TraceableSpeech: Towards Proactively Traceable Text-to-Speech with Watermarking
    2024/06/07 by Zhou, Junzuo, Yi, Jiangyan, Wang, Tao +5 · 7 citations
    #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
  6. Gated Recurrent Fusion with Joint Training Framework for Robust End-to-End Speech Recognition
    2020/11/09 by Fan, Cunhang, Yi, Jiangyan, Tao, Jianhua +3 · 2 citations
    #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
  7. Fast End-to-End Speech Recognition via Non-Autoregressive Models and Cross-Modal Knowledge Transferring from BERT
    2021/02/15 by Ye Bai, Bai, Ye, Jiangyan Yi +9 · 2 citations
    Computer Science · #Speech Recognition and Synthesis #Music and Audio Processing #Natural Language Processing Techniques
  8. FSR: Accelerating the Inference Process of Transducer-Based Models by Applying Fast-Skip Regularization
    2021/04/07 by Zhengkun Tian, Tian, Zhengkun, Jiangyan Yi +9 · 2 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  9. Spike-Triggered Non-Autoregressive Transformer for End-to-End Speech Recognition
    2020/05/16 by Zhengkun Tian, Tian, Zhengkun, Jiangyan Yi +9 · 2 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #electronic engineering #information engineering
  10. TemplateRL: Structured Template-Guided Reinforcement Learning for LLM Reasoning
    2025/05/21 by Jinyang Wu, Chan-Yu Liao, Wu, Jinyang +13 · 9 citations
    Decision Sciences · #Complex Systems and Decision Making #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG)
  11. AStar: Boosting Multimodal Reasoning with Automated Structured Thinking
    2025/02/04 by Jinyang Wu, Mingkuan Feng, Wu, Jinyang +14 · 4 citations
    Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Natural Language Processing Techniques #Semantic Web and Ontologies #Speech and dialogue systems
  12. VQ-CTAP: Cross-Modal Fine-Grained Sequence Representation Learning for Speech Processing
    2024/08/11 by Chunyu Qiang, Geng Wang, Qiang, Chunyu +26 · 5 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and dialogue systems #electronic engineering #information engineering
  13. Towards Fine-Grained Prosody Control for Voice Conversion
    2019/10/24 by Zheng Lian, Zhengqi Wen, Lian, Zheng +1 · 1 citation
    Computer Science · #Speech Recognition and Synthesis #Music and Audio Processing #Speech and Audio Processing
  14. Learn Spelling from Teachers: Transferring Knowledge from Language Models to Sequence-to-Sequence Speech Recognition
    2019/07/13 by Ye Bai, Bai, Ye, Jiangyan Yi +7 · 1 citation
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Natural Language Processing Techniques #Speech Recognition and Synthesis #electronic engineering #information engineering
  15. Integrating Knowledge into End-to-End Speech Recognition from External Text-Only Data
    2019/12/04 by Ye Bai, Bai, Ye, Jiangyan Yi +9 · 1 citation
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Topic Modeling #electronic engineering #information engineering
  16. Listen Attentively, and Spell Once: Whole Sentence Generation via a Non-Autoregressive Architecture for Low-Latency Speech Recognition
    2020/05/11 by Ye Bai, Jiangyan Yi, Bai, Ye +9 · 1 citation
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Speech Recognition and Synthesis #Topic Modeling #electronic engineering #information engineering
  17. Generalized Source Tracing: Detecting Novel Audio Deepfake Algorithm with Real Emphasis and Fake Dispersion Strategy
    2024/06/05 by Yuankun Xie, Xie, Yuankun, Ruibo Fu +13 · 2 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Digital Media Forensic Detection #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech and Audio Processing #electronic engineering #information engineering
  18. Codecfake: An Initial Dataset for Detecting LLM-based Deepfake Audio
    2024/06/12 by Lu, Yi, Xie, Yuankun, Fu, Ruibo +9 · 2 citations
    #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
  19. Singing-Tacotron: Global duration control attention and dynamic filter for End-to-end singing voice synthesis
    2022/02/16 by Wang, Tao, Fu, Ruibo, Yi, Jiangyan +2 · 1 citation
    #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Sound (cs.SD) #electronic engineering #information engineering
  20. Exploring the Role of Audio in Multimodal Misinformation Detection
    2024/08/22 by Yukun Liu, Liu, Moyang, Liu, Yukun +10 · 2 citations
    Social Sciences · #Misinformation and Its Impacts
  21. MDPE: A Multimodal Deception Dataset with Personality and Emotional Characteristics
    2024/07/17 by Cong Cai, Cai, Cong, Shan Liang +25 · 2 citations
    Computer Science · Psychology · Social Sciences · #Cybercrime and Law Enforcement Studies #Deception detection and forensic psychology #Crime Patterns and Interventions
  22. Emotion Selectable End-to-End Text-based Speech Editing
    2022/12/20 by Wang, Tao, Yi, Jiangyan, Fu, Ruibo +3 · 1 citation
    #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Sound (cs.SD) #electronic engineering #information engineering
  23. Learning From Yourself: A Self-Distillation Method for Fake Speech Detection
    2023/03/02 by Cunhang Fan, Xue, Jun, Jiangyan Yi +10 · 1 citation
    Computer Science · #Audio and Speech Processing (eess.AS) #Digital Media Forensic Detection #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Multimedia (cs.MM) #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  24. ALLM4ADD: Unlocking the Capabilities of Audio Large Language Models for Audio Deepfake Detection
    2025/05/16 by Hao Gu, Jiangyan Yi, Gu, Hao +15 · 3 citations
    Computer Science · #Generative Adversarial Networks and Image Synthesis #Speech Recognition and Synthesis #Music and Audio Processing
  25. Genuine-Focused Learning using Mask AutoEncoder for Generalized Fake Audio Detection
    2024/06/05 by Wang, Xiaopeng, Fu, Ruibo, Wen, Zhengqi +9 · 2 citations
    #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
  26. DReSS: Data-driven Regularized Structured Streamlining for Large Language Models
    2025/01/29 by Mingkuan Feng, Feng, Mingkuan, Jinyang Wu +13 · 1 citation
    Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Natural Language Processing Techniques #Topic Modeling
  27. ASRRL-TTS: Agile Speaker Representation Reinforcement Learning for Text-to-Speech Speaker Adaptation
    2024/07/07 by Ruibo Fu, Xin Qi, Fu, Ruibo +23 · 1 citation
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  28. Two-Stage Regularization-Based Structured Pruning for LLMs
    2025/05/23 by Feng, Mingkuan, Wu, Jinyang, Liu, Siyuan +6 · 1 citation
    #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG)
  29. PPPR: Portable Plug-in Prompt Refiner for Text to Audio Generation
    2024/06/07 by Shi, Shuchen, Fu, Ruibo, Wen, Zhengqi +10 · 1 citation
    #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
  30. Does Current Deepfake Audio Detection Model Effectively Detect ALM-based Deepfake Audio?
    2024/08/20 by Yuankun Xie, Xie, Yuankun, Chenxu Xiong +21 · 1 citation
    Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Digital Media Forensic Detection #FOS: Computer and information sciences #FOS: Electrical engineering #Image and Signal Denoising Methods #Music and Audio Processing #Sound (cs.SD) #electronic engineering #information engineering
  31. M3ANet: Multi-scale and Multi-Modal Alignment Network for Brain-Assisted Target Speaker Extraction
    2025/05/31 by Cunhang Fan, Ying Chen, Fan, Cunhang +14 · 2 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  32. Neural Codec Source Tracing: Toward Comprehensive Attribution in Open-Set Condition
    2025/01/11 by Xie, Yuankun, Wang, Xiaopeng, Wang, Zhiyong +7 · 1 citation
    #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
  33. DPI-TTS: Directional Patch Interaction for Fast-Converging and Style Temporal Modeling in Text-to-Speech
    2024/09/18 by Qi, Xin, Fu, Ruibo, Wen, Zhengqi +12 · 1 citation
    #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
  34. ImViD: Immersive Volumetric Videos for Enhanced VR Engagement
    2025/03/18 by Yang, Zhengxian, Pan, Shi, Wang, Shengqi +7 · 1 citation
    #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences
  35. SecoustiCodec: Cross-Modal Aligned Streaming Single-Codecbook Speech Codec
    2025/08/04 by Qiang, Chunyu, Wang, Haoyu, Gong, Cheng +10 · 3 citations
    #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
  36. RadialRouter: Structured Representation for Efficient and Robust Large Language Models Routing
    2025/06/04 by Ruihan Jin, Pengpeng Shao, Jin, Ruihan +11 · 3 citations
    Computer Science · Physics and Astronomy · #Natural Language Processing Techniques #Complex Network Analysis Techniques #Advanced Graph Neural Networks
  37. P2Mark: Plug-and-play Parameter-level Watermarking for Neural Speech Generation
    2025/04/07 by Ren, Yong, Yi, Jiangyan, Wang, Tao +7 · 1 citation
    #FOS: Computer and information sciences #Sound (cs.SD)