Abdelaziz, Ahmed Hussen
- Modality Dropout for Improved Performance-driven Talking Faces
2020/05/27 by Abdelaziz, Ahmed Hussen, Theobald, Barry-John, Dixon, Paul +3 · 5 citations
#Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Sound (cs.SD) #electronic engineering #information engineering
- ESPnet-SPK: full pipeline speaker embedding toolkit with reproducible recipes, self-supervised front-ends, and off-the-shelf models
2024/01/30 by Jee-weon Jung, Wangyou Zhang, Jung, Jee-weon +13 · 9 citations
Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Topic Modeling #electronic engineering #information engineering
- Can you Remove the Downstream Model for Speaker Recognition with Self-Supervised Speech Features?
2024/02/01 by Zakaria Aldeneh, Aldeneh, Zakaria, Takuya Higuchi +15 · 3 citations
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #Speech Recognition and Synthesis #electronic engineering #information engineering
- Modality Dropout for Multimodal Device Directed Speech Detection using Verbal and Non-Verbal Features
2023/10/23 by Gautam Krishna, Sameer Dharur, Krishna, Gautam +11 · 1 citation
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Human-Computer Interaction (cs.HC) #Machine Learning (cs.LG) #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #Speech and dialogue systems #electronic engineering #information engineering
- Comparative Analysis of Personalized Voice Activity Detection Systems: Assessing Real-World Effectiveness
2024/06/12 by Kumar, Satyam, Buddi, Sai Srujana, Sarawgi, Utkarsh Oggy +6 · 1 citation
#Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Human-Computer Interaction (cs.HC) #Machine Learning (cs.LG) #electronic engineering #information engineering
- A Variational Framework for Improving Naturalness in Generative Spoken Language Models
2025/06/17 by Liwei Chen, Takuya Higuchi, Chen, Li-Wei +7 · 2 citations
Computer Science · #Topic Modeling #Natural Language Processing Techniques #Speech and dialogue systems