Yan, Zhiyong
- speechocean762: An Open-Source Non-native English Speech Corpus For Pronunciation Assessment
2021/04/03 by Junbo Zhang, Zhang, Junbo, Zhiwen Zhang +15 · 13 citations
Computer Science · Psychology · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Phonetics and Phonology Research #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and dialogue systems #electronic engineering #information engineering
- Scaling up masked audio encoder learning for general audio classification
2024/06/11 by Heinrich Dinkel, Zhiyong Yan, Dinkel, Heinrich +9 · 24 citations
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech and Audio Processing #electronic engineering #information engineering
- CED: Consistent ensemble distillation for audio tagging
2023/08/23 by Heinrich Dinkel, Dinkel, Heinrich, Yongqing Wang +7 · 11 citations
Arts and Humanities · Computer Science · #Audio and Speech Processing (eess.AS) #Diverse Musicological Studies #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech and Audio Processing #electronic engineering #information engineering
- AV-SepFormer: Cross-Attention SepFormer for Audio-Visual Target Speaker Extraction
2023/06/25 by Jiuxin Lin, Lin, Jiuxin, Xinyu Cai +17 · 8 citations
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Multimedia (cs.MM) #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- Bridging Language Gaps in Audio-Text Retrieval
2024/06/11 by Yan, Zhiyong, Dinkel, Heinrich, Wang, Yongqing +4 · 2 citations
#Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
- Pseudo strong labels for large scale weakly supervised audio tagging
2022/04/28 by Dinkel, Heinrich, Yan, Zhiyong, Wang, Yongqing +2 · 1 citation
#Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
- Enhancing Automated Audio Captioning via Large Language Models with Optimized Audio Encoding
2024/06/19 by Liu, Jizhong, Li, Gang, Zhang, Junbo +5 · 2 citations
#Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
- GLAP: General contrastive audio-text pretraining across domains and languages
2025/06/12 by Dinkel, Heinrich, Yan, Zhiyong, Wang, Tianzi +7 · 4 citations
#Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
- Focus on the Sound around You: Monaural Target Speaker Extraction via Distance and Speaker Information
2023/06/28 by Lin, Jiuxin, Wang, Peng, Dinkel, Heinrich +6 · 1 citation
#Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering