Zhiyao Duan
- Deep Cross-Modal Audio-Visual Generation
2017/04/26 by Lele Chen, Sudhanshu Srivastava, Chen, Lele +5 · 1 voice · 5 citations
Computer Science · #Music and Audio Processing #Music Technology and Sound Studies #Generative Adversarial Networks and Image Synthesis
- Audio-Visual Event Localization in Unconstrained Videos
2018/03/23 by Yapeng Tian, Tian, Yapeng, Jing Shi +7 · 58 citations
Computer Science · #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Music and Audio Processing #Speech and Audio Processing #Video Analysis and Summarization
- Lip Movements Generation at a Glance
2018/03/28 by Lele Chen, Chen, Lele, Zhiheng Li +7 · 17 citations
Computer Science · #Speech and Audio Processing #Face recognition and analysis #Music and Audio Processing
- Automatic Music Transcription: An Overview
2018/12/25 by Emmanouil Benetos, Simon Dixon, Zhiyao Duan +1 · 22 citations
Computer Science · #Music and Audio Processing #Speech and Audio Processing #Music Technology and Sound Studies
- SVDD Challenge 2024: A Singing Voice Deepfake Detection Challenge (CtrSVDD Track, Test Set)
2023/09/14 by You Zhang, Yongyi Zang, Zang, Yongyi +12 · 13 citations
Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- UR Channel-Robust Synthetic Speech Detection System for ASVspoof 2021
2021/07/26 by Xinhui Chen, You Zhang, Chen, Xinhui +5 · 7 citations
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- Scoring Time Intervals using Non-Hierarchical Transformer For Automatic Piano Transcription
2024/04/15 by Yujia Yan, Yan, Yujia, Zhiyao Duan +1 · 12 citations
Arts and Humanities · Computer Science · #Audio and Speech Processing (eess.AS) #Diverse Musicological Studies #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Music Technology and Sound Studies #Music and Audio Processing #Sound (cs.SD) #electronic engineering #information engineering
- RL-Duet: Online Music Accompaniment Generation Using Deep Reinforcement Learning
2020/02/08 by Nan Jiang, Jiang, Nan, Sheng Jin +5 · 4 citations
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Human-Computer Interaction (cs.HC) #Machine Learning (cs.LG) #Music Technology and Sound Studies #Music and Audio Processing #Reinforcement Learning in Robotics #electronic engineering #information engineering
- Speech Driven Talking Face Generation from a Single Image and an Emotion Condition
2020/08/08 by Şefik Emre Eskimez, You Zhang, Eskimez, Sefik Emre +3 · 4 citations
Computer Science · #Audio and Speech Processing (eess.AS) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #FOS: Electrical engineering #Face recognition and analysis #Generative Adversarial Networks and Image Synthesis #Machine Learning (cs.LG) #Multimedia (cs.MM) #Speech and Audio Processing #electronic engineering #information engineering
- Rethinking Audio-visual Synchronization for Active Speaker Detection
2022/06/21 by Abudukelimu Wuerkaixi, Wuerkaixi, Abudukelimu, You Zhang +5 · 4 citations
Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #FOS: Electrical engineering #Multimedia (cs.MM) #Music and Audio Processing #Sound (cs.SD) #Speech and Audio Processing #Video Analysis and Summarization #electronic engineering #information engineering
- SynthTab: Leveraging Synthesized Data for Guitar Tablature Transcription
2023/09/16 by Yongyi Zang, Yi Zhong, Zang, Yongyi +5 · 5 citations
Arts and Humanities · Computer Science · #Audio and Speech Processing (eess.AS) #Diverse Musicological Studies #FOS: Computer and information sciences #FOS: Electrical engineering #Information Retrieval (cs.IR) #Multimedia (cs.MM) #Music Technology and Sound Studies #Music and Audio Processing #Signal Processing (eess.SP) #Sound (cs.SD) #electronic engineering #information engineering
- An Empirical Study on Channel Effects for Synthetic Voice Spoofing Countermeasure Systems
2021/04/03 by You Zhang, Ge Zhu, Zhang, You +5 · 4 citations
Computer Science · #Speech Recognition and Synthesis #Speech and Audio Processing #Natural Language Processing Techniques
- SAMO: Speaker Attractor Multi-Center One-Class Learning for Voice Anti-Spoofing
2022/11/04 by Siwen Ding, Ding, Siwen, You Zhang +3 · 4 citations
Computer Science · Medicine · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Signal Processing (eess.SP) #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #Voice and Speech Disorders #electronic engineering #information engineering
- BeatNet: CRNN and Particle Filtering for Online Joint Beat Downbeat and Meter Tracking
2021/08/08 by Mojtaba Heydari, Frank Cwitkowitz, Heydari, Mojtaba +3 · 3 citations
Computer Science · Neuroscience · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Information Retrieval (cs.IR) #Machine Learning (cs.LG) #Music and Audio Processing #Neuroscience and Music Perception #Signal Processing (eess.SP) #Sound (cs.SD) #Speech and Audio Processing #electronic engineering #information engineering
- SVDD 2024: The Inaugural Singing Voice Deepfake Detection Challenge
2024/08/28 by You Zhang, Zhang, You, Yongyi Zang +9 · 7 citations
Computer Science · #Music and Audio Processing #Speech and Audio Processing #Speech Recognition and Synthesis
- A Probabilistic Fusion Framework for Spoofing Aware Speaker Verification
2022/02/10 by You Zhang, Ge Zhu, Zhang, You +3 · 3 citations
Computer Science · Medicine · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #Voice and Speech Disorders #electronic engineering #information engineering
- Singing Beat Tracking With Self-supervised Front-end and Linear Transformers
2022/08/31 by Mojtaba Heydari, Heydari, Mojtaba, Zhiyao Duan +1 · 3 citations
Computer Science · #Music and Audio Processing #Speech and Audio Processing #Music Technology and Sound Studies
- EDMSound: Spectrogram Based Diffusion Models for Efficient and High-Quality Audio Synthesis
2023/11/15 by Ge Zhu, Yutong Wen, Zhu, Ge +5 · 4 citations
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music Technology and Sound Studies #Music and Audio Processing #Sound (cs.SD) #Speech and Audio Processing #electronic engineering #information engineering
- MusicHiFi: Fast High-Fidelity Stereo Vocoding
2024/03/15 by Ge Zhu, Juan-Pablo Cáceres, Zhu, Ge +5 · 3 citations
Computer Science · Engineering · #Advanced Adaptive Filtering Techniques #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music Technology and Sound Studies #Signal Processing (eess.SP) #Sound (cs.SD) #Speech and Audio Processing #electronic engineering #information engineering
- Emotion Classification: How Does an Automated System Compare to Naive Human Coders?
2015/10/22 by Sefik Emre Eskimez, Şefik Emre Eskimez, Eskimez, Sefik Emre +11 · 1 citation
Computer Science · Psychology · #Emotion and Mood Recognition #Sentiment Analysis and Opinion Mining #cs.HC
- Mitigating Cross-Database Differences for Learning Unified HRTF Representation
2023/07/27 by Yutong Wen, Wen, Yutong, You Zhang +3 · 2 citations
Computer Science · Earth and Planetary Sciences · Neuroscience · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Hearing Loss and Rehabilitation #Sound (cs.SD) #Speech and Audio Processing #Underwater Acoustics Research #electronic engineering #information engineering
- GTR-Voice: Articulatory Phonetics Informed Controllable Expressive Speech Synthesis
2024/06/15 by Zehua Kcriss Li, Meiying Melissa Chen, Li, Zehua Kcriss +7 · 3 citations
Psychology · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Phonetics and Phonology Research #Sound (cs.SD) #electronic engineering #information engineering
- Audiovisual Singing Voice Separation
2021/07/01 by Bochen Li, Yuxuan Wang, Li, Bochen +3 · 1 citation
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Music Technology and Sound Studies #Music and Audio Processing #Sound (cs.SD) #Speech and Audio Processing #electronic engineering #information engineering
- A Multi-Stream Fusion Approach with One-Class Learning for Audio-Visual Deepfake Detection
2024/06/20 by Kyung‐Bok Lee, You Zhang, Lee, Kyungbok +3 · 2 citations
Computer Science · #Anomaly Detection Techniques and Applications #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Digital Media Forensic Detection #FOS: Computer and information sciences #FOS: Electrical engineering #Multimedia (cs.MM) #Sound (cs.SD) #Speech and Audio Processing #electronic engineering #information engineering
- Generating Data with Text-to-Speech and Large-Language Models for Conversational Speech Recognition
2024/08/17 by Samuele Cornell, Cornell, Samuele, Jordan Darefsky +5 · 2 citations
Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #Speech Recognition and Synthesis #electronic engineering #information engineering
- Learning Arousal-Valence Representation from Categorical Emotion Labels of Speech
2023/11/24 by Enting Zhou, Zhou, Enting, You Zhang +3 · 1 citation
Computer Science · Psychology · #Audio and Speech Processing (eess.AS) #Emotion and Mood Recognition #FOS: Electrical engineering #Sentiment Analysis and Opinion Mining #Speech Recognition and Synthesis #electronic engineering #information engineering
- Toward Fully Self-Supervised Multi-Pitch Estimation
2024/02/23 by Frank Cwitkowitz, Zhiyao Duan, Cwitkowitz, Frank +1 · 2 citations
Engineering · #Advancements in Photolithography Techniques
- Audio Visual Segmentation Through Text Embeddings
2025/02/22 by Kyung‐Bok Lee, You Zhang, Lee, Kyungbok +3 · 1 citation
Computer Science · #Music and Audio Processing #Video Analysis and Summarization #Speech Recognition and Synthesis
- SVDD Challenge 2024: A Singing Voice Deepfake Detection Challenge Evaluation Plan
2024/05/08 by You Zhang, Zhang, You, Yongyi Zang +13 · 1 citation
Computer Science · #Music and Audio Processing #Speech Recognition and Synthesis #Speech and Audio Processing
- PartialEdit: Identifying Partial Deepfakes in the Era of Neural Speech Editing
2025/06/03 by You Zhang, Zhang, You, Linbo Zhang +4 · 2 citations
Computer Science · #Topic Modeling #Speech Recognition and Synthesis
- Conan: A Chunkwise Online Network for Zero-Shot Adaptive Voice Conversion
2025/07/19 by Yu Zhang, Zhang, Yu, Tian, Baotong +2 · 1 citation
Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- Summary of DCASE 2026 Task 5: Audio-Dependent Question Answering
2026/07/21 by Haolin He, Renhe Sun, Zheqi Dai +16
Engineering · #eess.AS