Hsuan-Fu Wang
- SpeechCLIP+: Self-supervised multi-task representation learning for speech via CLIP and speech-image data
2024/02/10 by Hsuan-Fu Wang, Wang, Hsuan-Fu, Yi-Jen Shih +13 · 2 citations
Computer Science · Psychology · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Phonetics and Phonology Research #Sound (cs.SD) #Speech Recognition and Synthesis #electronic engineering #information engineering
- M-SpeechCLIP: Leveraging Large-Scale, Pre-Trained Models for Multilingual Speech to Image Retrieval
2022/11/02 by Layne Berry, Yi-Jen Shih, Berry, Layne +9 · 2 citations
Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #Domain Adaptation and Few-Shot Learning #FOS: Computer and information sciences #FOS: Electrical engineering #Multimodal Machine Learning Applications #Sound (cs.SD) #Topic Modeling #electronic engineering #information engineering