Katrin Kirchhoff
- SpeechVerse: A Large-scale Generalizable Audio Language Model
2024/05/14 by Nilaksh Das, Das, Nilaksh, Saket Dingliwal +29 · 24 citations
Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- AutoGluon-Multimodal (AutoMM): Supercharging Multimodal AutoML with Foundation Models
2024/04/24 by Zhiqiang Tang, Haoyang Fang, Tang, Zhiqiang +13 · 12 citations
Computer Science · #Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Natural Language Processing Techniques
- Align-Refine: Non-Autoregressive Speech Recognition via Iterative Realignment
2020/10/24 by Ethan A. Chi, Julian Salazar, Chi, Ethan A. +3 · 4 citations
Computer Science · Engineering · #Music and Audio Processing #Speech Recognition and Synthesis #Speech and Audio Processing #cs.CL #cs.LG #cs.SD #eess.AS
- CriSPO: Multi-Aspect Critique-Suggestion-guided Automatic Prompt Optimization for Text Generation
2024/10/03 by He Han, He, Han, Qianchu Liu +11 · 7 citations
Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Natural Language Processing Techniques #Software Engineering Research #Topic Modeling
- Masked Language Model Scoring
2019/10/31 by Julian Salazar, Davis Liang, Toan Q. Nguyen +1 · 2 citations
Computer Science · Engineering · Mathematics · #cs.CL #cs.LG #eess.AS #stat.ML
- SpeechGuard: Exploring the Adversarial Robustness of Multimodal Large Language Models
2024/05/14 by Raghuveer Peri, Sai Muralidhar Jayanthi, Peri, Raghuveer +26 · 1 voice · 3 citations
Computer Science · Engineering · #Adversarial Robustness in Machine Learning #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #cs.CL #cs.SD #eess.AS #electronic engineering #information engineering
- Zero-resource Speech Translation and Recognition with LLMs
2024/12/24 by Karel Mundnich, Xing Niu, Mundnich, Karel +23 · 5 citations
Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #electronic engineering #information engineering
- Multimodal Semi-supervised Learning Framework for Punctuation Prediction in Conversational Speech
2020/08/03 by Monica Sunkara, Srikanth Ronanki, Sunkara, Monica +7 · 1 citation
Computer Science · Engineering · #Natural Language Processing Techniques #Speech Recognition and Synthesis #Speech and dialogue systems #cs.CL #eess.AS
- Transformer-Transducers for Code-Switched Speech Recognition
2020/11/30 by Siddharth Dalmia, Dalmia, Siddharth, Yuzong Liu +5 · 1 citation
Computer Science · Engineering · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #cs.CL #eess.AS #electronic engineering #information engineering
- Neural Inverse Text Normalization
2021/02/12 by Monica Sunkara, Chaitanya Shivade, Sunkara, Monica +5 · 1 citation
Computer Science · Engineering · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #cs.CL #eess.AS #electronic engineering #information engineering
- Representation learning through cross-modal conditional teacher-student training for speech emotion recognition
2021/11/30 by Sundararajan Srinivasan, Zhaocheng Huang, Srinivasan, Sundararajan +3 · 1 citation
Computer Science · Engineering · Psychology · #Emotion and Mood Recognition #Sentiment Analysis and Opinion Mining #Speech Recognition and Synthesis #eess.AS
- Mask The Bias: Improving Domain-Adaptive Generalization of CTC-based ASR with Internal Language Model Estimation
2023/05/05 by Nilaksh Das, Das, Nilaksh, Monica Sunkara +11 · 1 citation
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Topic Modeling #electronic engineering #information engineering
- Cross-dialectal data sharing for acoustic modeling in Arabic speech recognition
2005/03/12 by Katrin Kirchhoff, Dimitra Vergyri · 1 citation
Computer Science · #Acoustic model #Arabic #Artificial intelligence #Computer science #Language model #Linguistics #Modern Standard Arabic #Natural Language Processing Techniques #Natural language processing #Phonetic transcription #Speech Recognition and Synthesis #Speech corpus #Speech processing #Speech recognition #Speech synthesis #Topic Modeling #Transcription (linguistics) #Word error rate