vix.ing · top · new · best · stats · spec

Zhiyao Duan

  1. Deep Cross-Modal Audio-Visual Generation
    2017/04/26 by Lele Chen, Sudhanshu Srivastava, Chen, Lele +5 · 1 voice · 5 citations
    Computer Science · #Music and Audio Processing #Music Technology and Sound Studies #Generative Adversarial Networks and Image Synthesis
  2. Audio-Visual Event Localization in Unconstrained Videos
    2018/03/23 by Yapeng Tian, Tian, Yapeng, Jing Shi +7 · 58 citations
    Computer Science · #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Music and Audio Processing #Speech and Audio Processing #Video Analysis and Summarization
  3. Lip Movements Generation at a Glance
    2018/03/28 by Lele Chen, Chen, Lele, Zhiheng Li +7 · 17 citations
    Computer Science · #Speech and Audio Processing #Face recognition and analysis #Music and Audio Processing
  4. Automatic Music Transcription: An Overview
    2018/12/25 by Emmanouil Benetos, Simon Dixon, Zhiyao Duan +1 · 22 citations
    Computer Science · #Music and Audio Processing #Speech and Audio Processing #Music Technology and Sound Studies
  5. SVDD Challenge 2024: A Singing Voice Deepfake Detection Challenge (CtrSVDD Track, Test Set)
    2023/09/14 by You Zhang, Yongyi Zang, Zang, Yongyi +12 · 13 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  6. UR Channel-Robust Synthetic Speech Detection System for ASVspoof 2021
    2021/07/26 by Xinhui Chen, You Zhang, Chen, Xinhui +5 · 7 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  7. Scoring Time Intervals using Non-Hierarchical Transformer For Automatic Piano Transcription
    2024/04/15 by Yujia Yan, Yan, Yujia, Zhiyao Duan +1 · 12 citations
    Arts and Humanities · Computer Science · #Audio and Speech Processing (eess.AS) #Diverse Musicological Studies #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Music Technology and Sound Studies #Music and Audio Processing #Sound (cs.SD) #electronic engineering #information engineering
  8. RL-Duet: Online Music Accompaniment Generation Using Deep Reinforcement Learning
    2020/02/08 by Nan Jiang, Jiang, Nan, Sheng Jin +5 · 4 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Human-Computer Interaction (cs.HC) #Machine Learning (cs.LG) #Music Technology and Sound Studies #Music and Audio Processing #Reinforcement Learning in Robotics #electronic engineering #information engineering
  9. Speech Driven Talking Face Generation from a Single Image and an Emotion Condition
    2020/08/08 by Şefik Emre Eskimez, You Zhang, Eskimez, Sefik Emre +3 · 4 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #FOS: Electrical engineering #Face recognition and analysis #Generative Adversarial Networks and Image Synthesis #Machine Learning (cs.LG) #Multimedia (cs.MM) #Speech and Audio Processing #electronic engineering #information engineering
  10. Rethinking Audio-visual Synchronization for Active Speaker Detection
    2022/06/21 by Abudukelimu Wuerkaixi, Wuerkaixi, Abudukelimu, You Zhang +5 · 4 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #FOS: Electrical engineering #Multimedia (cs.MM) #Music and Audio Processing #Sound (cs.SD) #Speech and Audio Processing #Video Analysis and Summarization #electronic engineering #information engineering
  11. SynthTab: Leveraging Synthesized Data for Guitar Tablature Transcription
    2023/09/16 by Yongyi Zang, Yi Zhong, Zang, Yongyi +5 · 5 citations
    Arts and Humanities · Computer Science · #Audio and Speech Processing (eess.AS) #Diverse Musicological Studies #FOS: Computer and information sciences #FOS: Electrical engineering #Information Retrieval (cs.IR) #Multimedia (cs.MM) #Music Technology and Sound Studies #Music and Audio Processing #Signal Processing (eess.SP) #Sound (cs.SD) #electronic engineering #information engineering
  12. An Empirical Study on Channel Effects for Synthetic Voice Spoofing Countermeasure Systems
    2021/04/03 by You Zhang, Ge Zhu, Zhang, You +5 · 4 citations
    Computer Science · #Speech Recognition and Synthesis #Speech and Audio Processing #Natural Language Processing Techniques
  13. SAMO: Speaker Attractor Multi-Center One-Class Learning for Voice Anti-Spoofing
    2022/11/04 by Siwen Ding, Ding, Siwen, You Zhang +3 · 4 citations
    Computer Science · Medicine · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Signal Processing (eess.SP) #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #Voice and Speech Disorders #electronic engineering #information engineering
  14. BeatNet: CRNN and Particle Filtering for Online Joint Beat Downbeat and Meter Tracking
    2021/08/08 by Mojtaba Heydari, Frank Cwitkowitz, Heydari, Mojtaba +3 · 3 citations
    Computer Science · Neuroscience · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Information Retrieval (cs.IR) #Machine Learning (cs.LG) #Music and Audio Processing #Neuroscience and Music Perception #Signal Processing (eess.SP) #Sound (cs.SD) #Speech and Audio Processing #electronic engineering #information engineering
  15. SVDD 2024: The Inaugural Singing Voice Deepfake Detection Challenge
    2024/08/28 by You Zhang, Zhang, You, Yongyi Zang +9 · 7 citations
    Computer Science · #Music and Audio Processing #Speech and Audio Processing #Speech Recognition and Synthesis
  16. A Probabilistic Fusion Framework for Spoofing Aware Speaker Verification
    2022/02/10 by You Zhang, Ge Zhu, Zhang, You +3 · 3 citations
    Computer Science · Medicine · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #Voice and Speech Disorders #electronic engineering #information engineering
  17. Singing Beat Tracking With Self-supervised Front-end and Linear Transformers
    2022/08/31 by Mojtaba Heydari, Heydari, Mojtaba, Zhiyao Duan +1 · 3 citations
    Computer Science · #Music and Audio Processing #Speech and Audio Processing #Music Technology and Sound Studies
  18. EDMSound: Spectrogram Based Diffusion Models for Efficient and High-Quality Audio Synthesis
    2023/11/15 by Ge Zhu, Yutong Wen, Zhu, Ge +5 · 4 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music Technology and Sound Studies #Music and Audio Processing #Sound (cs.SD) #Speech and Audio Processing #electronic engineering #information engineering
  19. MusicHiFi: Fast High-Fidelity Stereo Vocoding
    2024/03/15 by Ge Zhu, Juan-Pablo Cáceres, Zhu, Ge +5 · 3 citations
    Computer Science · Engineering · #Advanced Adaptive Filtering Techniques #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music Technology and Sound Studies #Signal Processing (eess.SP) #Sound (cs.SD) #Speech and Audio Processing #electronic engineering #information engineering
  20. Emotion Classification: How Does an Automated System Compare to Naive Human Coders?
    2015/10/22 by Sefik Emre Eskimez, Şefik Emre Eskimez, Eskimez, Sefik Emre +11 · 1 citation
    Computer Science · Psychology · #Emotion and Mood Recognition #Sentiment Analysis and Opinion Mining #cs.HC
  21. Mitigating Cross-Database Differences for Learning Unified HRTF Representation
    2023/07/27 by Yutong Wen, Wen, Yutong, You Zhang +3 · 2 citations
    Computer Science · Earth and Planetary Sciences · Neuroscience · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Hearing Loss and Rehabilitation #Sound (cs.SD) #Speech and Audio Processing #Underwater Acoustics Research #electronic engineering #information engineering
  22. GTR-Voice: Articulatory Phonetics Informed Controllable Expressive Speech Synthesis
    2024/06/15 by Zehua Kcriss Li, Meiying Melissa Chen, Li, Zehua Kcriss +7 · 3 citations
    Psychology · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Phonetics and Phonology Research #Sound (cs.SD) #electronic engineering #information engineering
  23. Audiovisual Singing Voice Separation
    2021/07/01 by Bochen Li, Yuxuan Wang, Li, Bochen +3 · 1 citation
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Music Technology and Sound Studies #Music and Audio Processing #Sound (cs.SD) #Speech and Audio Processing #electronic engineering #information engineering
  24. A Multi-Stream Fusion Approach with One-Class Learning for Audio-Visual Deepfake Detection
    2024/06/20 by Kyung‐Bok Lee, You Zhang, Lee, Kyungbok +3 · 2 citations
    Computer Science · #Anomaly Detection Techniques and Applications #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Digital Media Forensic Detection #FOS: Computer and information sciences #FOS: Electrical engineering #Multimedia (cs.MM) #Sound (cs.SD) #Speech and Audio Processing #electronic engineering #information engineering
  25. Generating Data with Text-to-Speech and Large-Language Models for Conversational Speech Recognition
    2024/08/17 by Samuele Cornell, Cornell, Samuele, Jordan Darefsky +5 · 2 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #Speech Recognition and Synthesis #electronic engineering #information engineering
  26. Learning Arousal-Valence Representation from Categorical Emotion Labels of Speech
    2023/11/24 by Enting Zhou, Zhou, Enting, You Zhang +3 · 1 citation
    Computer Science · Psychology · #Audio and Speech Processing (eess.AS) #Emotion and Mood Recognition #FOS: Electrical engineering #Sentiment Analysis and Opinion Mining #Speech Recognition and Synthesis #electronic engineering #information engineering
  27. Toward Fully Self-Supervised Multi-Pitch Estimation
    2024/02/23 by Frank Cwitkowitz, Zhiyao Duan, Cwitkowitz, Frank +1 · 2 citations
    Engineering · #Advancements in Photolithography Techniques
  28. Audio Visual Segmentation Through Text Embeddings
    2025/02/22 by Kyung‐Bok Lee, You Zhang, Lee, Kyungbok +3 · 1 citation
    Computer Science · #Music and Audio Processing #Video Analysis and Summarization #Speech Recognition and Synthesis
  29. SVDD Challenge 2024: A Singing Voice Deepfake Detection Challenge Evaluation Plan
    2024/05/08 by You Zhang, Zhang, You, Yongyi Zang +13 · 1 citation
    Computer Science · #Music and Audio Processing #Speech Recognition and Synthesis #Speech and Audio Processing
  30. PartialEdit: Identifying Partial Deepfakes in the Era of Neural Speech Editing
    2025/06/03 by You Zhang, Zhang, You, Linbo Zhang +4 · 2 citations
    Computer Science · #Topic Modeling #Speech Recognition and Synthesis
  31. Conan: A Chunkwise Online Network for Zero-Shot Adaptive Voice Conversion
    2025/07/19 by Yu Zhang, Zhang, Yu, Tian, Baotong +2 · 1 citation
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  32. Summary of DCASE 2026 Task 5: Audio-Dependent Question Answering
    2026/07/21 by Haolin He, Renhe Sun, Zheqi Dai +16
    Engineering · #eess.AS