vix.ing · top · new · best · stats · spec

Balam, Jagadeesh

  1. Fast Conformer with Linearly Scalable Attention for Efficient Speech Recognition
    2023/05/08 by Rekesh, Dima, Koluguri, Nithin Rao, Kriman, Samuel +8 · 30 citations
    #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
  2. SPGISpeech: 5,000 hours of transcribed financial audio for fully\n formatted end-to-end speech recognition
    2021/04/05 by Patrick O’Neill, Vitaly Lavrukhin, O'Neill, Patrick K. +24 · 12 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Natural Language Processing Techniques #Speech Recognition and Synthesis #electronic engineering #information engineering
  3. SALM: Speech-augmented Language Model with In-context Learning for Speech Recognition and Translation
    2023/10/13 by Zhehuai Chen, He Huang, Chen, Zhehuai +15 · 15 citations
    Computer Science · #68T10 #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Human-Computer Interaction (cs.HC) #I.2.7 #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and dialogue systems #electronic engineering #information engineering
  4. Schrödinger Bridge for Generative Speech Enhancement
    2024/07/22 by Jukić, Ante, Korostik, Roman, Balam, Jagadeesh +1 · 14 citations
    #Audio and Speech Processing (eess.AS) #FOS: Electrical engineering #electronic engineering #information engineering
  5. DeSTA2: Developing Instruction-Following Speech Language Model Without Speech Instruction-Tuning Data
    2024/09/30 by Ke-Han Lu, Zhehuai Chen, Lu, Ke-Han +13 · 15 citations
    Computer Science · #Speech and dialogue systems #Speech Recognition and Synthesis
  6. Sortformer: A Novel Approach for Permutation-Resolved Speaker Supervision in Speech-to-Text Systems
    2024/09/10 by Park, Taejin, Medennikov, Ivan, Dhawan, Kunal +6 · 12 citations
    #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Sound (cs.SD) #electronic engineering #information engineering
  7. Chain-of-Thought Prompting for Speech Translation
    2024/09/17 by Ke Hu, Hu, Ke, Zhehuai Chen +13 · 6 citations
    Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Natural Language Processing Techniques
  8. SALM-Duplex: Efficient and Direct Duplex Modeling for Speech-to-Speech Language Model
    2025/05/21 by Ke Hu, Ehsan Hosseini-Asl, Hu, Ke +17 · 11 citations
    Computer Science · #Speech Recognition and Synthesis #Speech and dialogue systems #Speech and Audio Processing
  9. BESTOW: Efficient and Streamable Speech Language Model with the Best of Two Worlds in GPT and T5
    2024/06/28 by Zhehuai Chen, He Huang, Chen, Zhehuai +13 · 5 citations
    Computer Science · #68T10 #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Human-Computer Interaction (cs.HC) #I.2.7 #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Topic Modeling #electronic engineering #information engineering
  10. Enhancing Speaker Diarization with Large Language Models: A Contextual Beam Search Approach
    2023/09/11 by Tae Jin Park, Park, Tae Jin, Kunal Dhawan +5 · 3 citations
    Computer Science · #Speech Recognition and Synthesis #Topic Modeling #Music and Audio Processing
  11. Genetic Instruct: Scaling up Synthetic Generation of Coding Instructions for Large Language Models
    2024/07/29 by Majumdar, Somshubra, Noroozi, Vahid, Samadi, Mehrzad +6 · 6 citations
    #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Neural and Evolutionary Computing (cs.NE)
  12. Discrete Audio Representation as an Alternative to Mel-Spectrograms for Speaker and Speech Recognition
    2023/09/19 by Puvvada, Krishna C., Koluguri, Nithin Rao, Dhawan, Kunal +2 · 3 citations
    #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
  13. Multi-scale Speaker Diarization with Dynamic Scale Weighting
    2022/03/30 by Park, Tae Jin, Koluguri, Nithin Rao, Balam, Jagadeesh +1 · 2 citations
    #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #electronic engineering #information engineering
  14. Training and Inference Efficiency of Encoder-Decoder Speech Models
    2025/03/07 by Żelasko, Piotr, Dhawan, Kunal, Galvez, Daniel +7 · 6 citations
    #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #electronic engineering #information engineering
  15. Investigating End-to-End ASR Architectures for Long Form Audio Transcription
    2023/09/18 by Nithin Rao Koluguri, Samuel Kriman, Koluguri, Nithin Rao +13 · 2 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  16. Granary: Speech Recognition and Translation Dataset in 25 European Languages
    2025/05/19 by Koluguri, Nithin Rao, Sekoyan, Monica, Zelenfroynd, George +12 · 6 citations
    #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #electronic engineering #information engineering
  17. Stateful Conformer with Cache-based Inference for Streaming Automatic Speech Recognition
    2023/12/27 by Noroozi, Vahid, Majumdar, Somshubra, Kumar, Ankur +2 · 2 citations
    #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #electronic engineering #information engineering
  18. Citrinet: Closing the Gap between Non-Autoregressive and Autoregressive End-to-End Models for Automatic Speech Recognition
    2021/04/05 by Majumdar, Somshubra, Balam, Jagadeesh, Hrinchuk, Oleksii +3 · 1 citation
    #Audio and Speech Processing (eess.AS) #FOS: Electrical engineering #electronic engineering #information engineering
  19. META-CAT: Speaker-Informed Speech Embeddings via Meta Information Concatenation for Multi-talker ASR
    2024/09/18 by Jinhan Wang, Wang, Jinhan, Weiqing Wang +17 · 2 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #Speech and dialogue systems #electronic engineering #information engineering
  20. A Compact End-to-End Model with Local and Global Context for Spoken Language Identification
    2022/10/27 by Fei Jia, Jia, Fei, Nithin Rao Koluguri +5 · 1 citation
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  21. Leveraging Pretrained ASR Encoders for Effective and Efficient End-to-End Speech Intent Classification and Slot Filling
    2023/07/13 by Huang, He, Balam, Jagadeesh, Ginsburg, Boris · 1 citation
    #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
  22. Canary-1B-v2 & Parakeet-TDT-0.6B-v3: Efficient and High-Performance Models for Multilingual ASR and AST
    2025/09/17 by Monica Sekoyan, Sekoyan, Monica, Nithin Rao Koluguri +13 · 5 citations
    Computer Science · Psychology · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Phonetics and Phonology Research #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  23. NEST: Self-supervised Fast Conformer as All-purpose Seasoning to Speech Processing Tasks
    2024/08/23 by Huang, He, Park, Taejin, Dhawan, Kunal +6 · 2 citations
    #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
  24. Multilingual Audio-Visual Speech Recognition with Hybrid CTC/RNN-T Fast Conformer
    2024/03/14 by Burchi, Maxime, Puvvada, Krishna C., Balam, Jagadeesh +2 · 1 citation
    #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #FOS: Electrical engineering #Multimedia (cs.MM) #Sound (cs.SD) #electronic engineering #information engineering
  25. Resource-Efficient Adaptation of Speech Foundation Models for Multi-Speaker ASR
    2024/09/02 by Wang, Weiqing, Dhawan, Kunal, Park, Taejin +6 · 1 citation
    #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
  26. NeKo: Cross-Modality Post-Recognition Error Correction with Tasks-Guided Mixture-of-Experts Language Model
    2024/11/08 by Yen‐Ting Lin, Zhehuai Chen, Lin, Yen-Ting +23 · 2 citations
    Computer Science · #Topic Modeling #Natural Language Processing Techniques
  27. EMMeTT: Efficient Multimodal Machine Translation Training
    2024/09/20 by Piotr Żelasko, Żelasko, Piotr, Zhehuai Chen +17 · 1 citation
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Sound (cs.SD) #Speech and dialogue systems #Topic Modeling #electronic engineering #information engineering
  28. SPGISpeech 2.0: Transcribed multi-speaker financial audio for speaker-tagged transcription
    2025/08/07 by Grossman, Raymond, Park, Taejin, Dhawan, Kunal +6 · 1 citation
    #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
  29. Large Language Model Based Generative Error Correction: A Challenge and Baselines for Speech Recognition, Speaker Tagging, and Emotion Recognition
    2024/09/15 by Chao-Han Huck Yang, Taejin Park, Yang, Chao-Han Huck +39 · 1 citation
    Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #electronic engineering #information engineering
  30. Longer is (Not Necessarily) Stronger: Punctuated Long-Sequence Training for Enhanced Speech Recognition and Translation
    2024/09/09 by Nithin Rao Koluguri, Koluguri, Nithin Rao, Travis Bartley +11 · 1 citation
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Speech Recognition and Synthesis #electronic engineering #information engineering
  31. VoiceTextBlender: Augmenting Large Language Models with Speech Capabilities via Single-Stage Joint Speech-Text Supervised Fine-Tuning
    2024/10/23 by Yifan Peng, Peng, Yifan, Krishna C. Puvvada +17 · 1 citation
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Speech Recognition and Synthesis #Topic Modeling #electronic engineering #information engineering
  32. Streaming Sortformer: Speaker Cache-Based Online Speaker Diarization with Arrival-Time Ordering
    2025/07/24 by Ivan Medennikov, Tae‐Jin Park, Medennikov, Ivan +13 · 3 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering