vix.ing · top · new · best · stats · spec

Katrin Kirchhoff

  1. SpeechVerse: A Large-scale Generalizable Audio Language Model
    2024/05/14 by Nilaksh Das, Das, Nilaksh, Saket Dingliwal +29 · 24 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  2. AutoGluon-Multimodal (AutoMM): Supercharging Multimodal AutoML with Foundation Models
    2024/04/24 by Zhiqiang Tang, Haoyang Fang, Tang, Zhiqiang +13 · 12 citations
    Computer Science · #Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Natural Language Processing Techniques
  3. Align-Refine: Non-Autoregressive Speech Recognition via Iterative Realignment
    2020/10/24 by Ethan A. Chi, Julian Salazar, Chi, Ethan A. +3 · 4 citations
    Computer Science · Engineering · #Music and Audio Processing #Speech Recognition and Synthesis #Speech and Audio Processing #cs.CL #cs.LG #cs.SD #eess.AS
  4. CriSPO: Multi-Aspect Critique-Suggestion-guided Automatic Prompt Optimization for Text Generation
    2024/10/03 by He Han, He, Han, Qianchu Liu +11 · 7 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Natural Language Processing Techniques #Software Engineering Research #Topic Modeling
  5. Masked Language Model Scoring
    2019/10/31 by Julian Salazar, Davis Liang, Toan Q. Nguyen +1 · 2 citations
    Computer Science · Engineering · Mathematics · #cs.CL #cs.LG #eess.AS #stat.ML
  6. SpeechGuard: Exploring the Adversarial Robustness of Multimodal Large Language Models
    2024/05/14 by Raghuveer Peri, Sai Muralidhar Jayanthi, Peri, Raghuveer +26 · 1 voice · 3 citations
    Computer Science · Engineering · #Adversarial Robustness in Machine Learning #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #cs.CL #cs.SD #eess.AS #electronic engineering #information engineering
  7. Zero-resource Speech Translation and Recognition with LLMs
    2024/12/24 by Karel Mundnich, Xing Niu, Mundnich, Karel +23 · 5 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #electronic engineering #information engineering
  8. Multimodal Semi-supervised Learning Framework for Punctuation Prediction in Conversational Speech
    2020/08/03 by Monica Sunkara, Srikanth Ronanki, Sunkara, Monica +7 · 1 citation
    Computer Science · Engineering · #Natural Language Processing Techniques #Speech Recognition and Synthesis #Speech and dialogue systems #cs.CL #eess.AS
  9. Transformer-Transducers for Code-Switched Speech Recognition
    2020/11/30 by Siddharth Dalmia, Dalmia, Siddharth, Yuzong Liu +5 · 1 citation
    Computer Science · Engineering · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #cs.CL #eess.AS #electronic engineering #information engineering
  10. Neural Inverse Text Normalization
    2021/02/12 by Monica Sunkara, Chaitanya Shivade, Sunkara, Monica +5 · 1 citation
    Computer Science · Engineering · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #cs.CL #eess.AS #electronic engineering #information engineering
  11. Representation learning through cross-modal conditional teacher-student training for speech emotion recognition
    2021/11/30 by Sundararajan Srinivasan, Zhaocheng Huang, Srinivasan, Sundararajan +3 · 1 citation
    Computer Science · Engineering · Psychology · #Emotion and Mood Recognition #Sentiment Analysis and Opinion Mining #Speech Recognition and Synthesis #eess.AS
  12. Mask The Bias: Improving Domain-Adaptive Generalization of CTC-based ASR with Internal Language Model Estimation
    2023/05/05 by Nilaksh Das, Das, Nilaksh, Monica Sunkara +11 · 1 citation
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Topic Modeling #electronic engineering #information engineering
  13. Cross-dialectal data sharing for acoustic modeling in Arabic speech recognition
    2005/03/12 by Katrin Kirchhoff, Dimitra Vergyri · 1 citation
    Computer Science · #Acoustic model #Arabic #Artificial intelligence #Computer science #Language model #Linguistics #Modern Standard Arabic #Natural Language Processing Techniques #Natural language processing #Phonetic transcription #Speech Recognition and Synthesis #Speech corpus #Speech processing #Speech recognition #Speech synthesis #Topic Modeling #Transcription (linguistics) #Word error rate