Balam, Jagadeesh
- Fast Conformer with Linearly Scalable Attention for Efficient Speech Recognition
2023/05/08 by Rekesh, Dima, Koluguri, Nithin Rao, Kriman, Samuel +8 · 30 citations
#Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
- SPGISpeech: 5,000 hours of transcribed financial audio for fully\n formatted end-to-end speech recognition
2021/04/05 by Patrick O’Neill, Vitaly Lavrukhin, O'Neill, Patrick K. +24 · 12 citations
Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Natural Language Processing Techniques #Speech Recognition and Synthesis #electronic engineering #information engineering
- SALM: Speech-augmented Language Model with In-context Learning for Speech Recognition and Translation
2023/10/13 by Zhehuai Chen, He Huang, Chen, Zhehuai +15 · 15 citations
Computer Science · #68T10 #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Human-Computer Interaction (cs.HC) #I.2.7 #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and dialogue systems #electronic engineering #information engineering
- Schrödinger Bridge for Generative Speech Enhancement
2024/07/22 by Jukić, Ante, Korostik, Roman, Balam, Jagadeesh +1 · 14 citations
#Audio and Speech Processing (eess.AS) #FOS: Electrical engineering #electronic engineering #information engineering
- DeSTA2: Developing Instruction-Following Speech Language Model Without Speech Instruction-Tuning Data
2024/09/30 by Ke-Han Lu, Zhehuai Chen, Lu, Ke-Han +13 · 15 citations
Computer Science · #Speech and dialogue systems #Speech Recognition and Synthesis
- Sortformer: A Novel Approach for Permutation-Resolved Speaker Supervision in Speech-to-Text Systems
2024/09/10 by Park, Taejin, Medennikov, Ivan, Dhawan, Kunal +6 · 12 citations
#Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Sound (cs.SD) #electronic engineering #information engineering
- Chain-of-Thought Prompting for Speech Translation
2024/09/17 by Ke Hu, Hu, Ke, Zhehuai Chen +13 · 6 citations
Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Natural Language Processing Techniques
- SALM-Duplex: Efficient and Direct Duplex Modeling for Speech-to-Speech Language Model
2025/05/21 by Ke Hu, Ehsan Hosseini-Asl, Hu, Ke +17 · 11 citations
Computer Science · #Speech Recognition and Synthesis #Speech and dialogue systems #Speech and Audio Processing
- BESTOW: Efficient and Streamable Speech Language Model with the Best of Two Worlds in GPT and T5
2024/06/28 by Zhehuai Chen, He Huang, Chen, Zhehuai +13 · 5 citations
Computer Science · #68T10 #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Human-Computer Interaction (cs.HC) #I.2.7 #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Topic Modeling #electronic engineering #information engineering
- Enhancing Speaker Diarization with Large Language Models: A Contextual Beam Search Approach
2023/09/11 by Tae Jin Park, Park, Tae Jin, Kunal Dhawan +5 · 3 citations
Computer Science · #Speech Recognition and Synthesis #Topic Modeling #Music and Audio Processing
- Genetic Instruct: Scaling up Synthetic Generation of Coding Instructions for Large Language Models
2024/07/29 by Majumdar, Somshubra, Noroozi, Vahid, Samadi, Mehrzad +6 · 6 citations
#Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Neural and Evolutionary Computing (cs.NE)
- Discrete Audio Representation as an Alternative to Mel-Spectrograms for Speaker and Speech Recognition
2023/09/19 by Puvvada, Krishna C., Koluguri, Nithin Rao, Dhawan, Kunal +2 · 3 citations
#Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
- Multi-scale Speaker Diarization with Dynamic Scale Weighting
2022/03/30 by Park, Tae Jin, Koluguri, Nithin Rao, Balam, Jagadeesh +1 · 2 citations
#Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #electronic engineering #information engineering
- Training and Inference Efficiency of Encoder-Decoder Speech Models
2025/03/07 by Żelasko, Piotr, Dhawan, Kunal, Galvez, Daniel +7 · 6 citations
#Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #electronic engineering #information engineering
- Investigating End-to-End ASR Architectures for Long Form Audio Transcription
2023/09/18 by Nithin Rao Koluguri, Samuel Kriman, Koluguri, Nithin Rao +13 · 2 citations
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- Granary: Speech Recognition and Translation Dataset in 25 European Languages
2025/05/19 by Koluguri, Nithin Rao, Sekoyan, Monica, Zelenfroynd, George +12 · 6 citations
#Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #electronic engineering #information engineering
- Stateful Conformer with Cache-based Inference for Streaming Automatic Speech Recognition
2023/12/27 by Noroozi, Vahid, Majumdar, Somshubra, Kumar, Ankur +2 · 2 citations
#Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #electronic engineering #information engineering
- Citrinet: Closing the Gap between Non-Autoregressive and Autoregressive End-to-End Models for Automatic Speech Recognition
2021/04/05 by Majumdar, Somshubra, Balam, Jagadeesh, Hrinchuk, Oleksii +3 · 1 citation
#Audio and Speech Processing (eess.AS) #FOS: Electrical engineering #electronic engineering #information engineering
- META-CAT: Speaker-Informed Speech Embeddings via Meta Information Concatenation for Multi-talker ASR
2024/09/18 by Jinhan Wang, Wang, Jinhan, Weiqing Wang +17 · 2 citations
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #Speech and dialogue systems #electronic engineering #information engineering
- A Compact End-to-End Model with Local and Global Context for Spoken Language Identification
2022/10/27 by Fei Jia, Jia, Fei, Nithin Rao Koluguri +5 · 1 citation
Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- Leveraging Pretrained ASR Encoders for Effective and Efficient End-to-End Speech Intent Classification and Slot Filling
2023/07/13 by Huang, He, Balam, Jagadeesh, Ginsburg, Boris · 1 citation
#Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
- Canary-1B-v2 & Parakeet-TDT-0.6B-v3: Efficient and High-Performance Models for Multilingual ASR and AST
2025/09/17 by Monica Sekoyan, Sekoyan, Monica, Nithin Rao Koluguri +13 · 5 citations
Computer Science · Psychology · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Phonetics and Phonology Research #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- NEST: Self-supervised Fast Conformer as All-purpose Seasoning to Speech Processing Tasks
2024/08/23 by Huang, He, Park, Taejin, Dhawan, Kunal +6 · 2 citations
#Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
- Multilingual Audio-Visual Speech Recognition with Hybrid CTC/RNN-T Fast Conformer
2024/03/14 by Burchi, Maxime, Puvvada, Krishna C., Balam, Jagadeesh +2 · 1 citation
#Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #FOS: Electrical engineering #Multimedia (cs.MM) #Sound (cs.SD) #electronic engineering #information engineering
- Resource-Efficient Adaptation of Speech Foundation Models for Multi-Speaker ASR
2024/09/02 by Wang, Weiqing, Dhawan, Kunal, Park, Taejin +6 · 1 citation
#Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
- NeKo: Cross-Modality Post-Recognition Error Correction with Tasks-Guided Mixture-of-Experts Language Model
2024/11/08 by Yen‐Ting Lin, Zhehuai Chen, Lin, Yen-Ting +23 · 2 citations
Computer Science · #Topic Modeling #Natural Language Processing Techniques
- EMMeTT: Efficient Multimodal Machine Translation Training
2024/09/20 by Piotr Żelasko, Żelasko, Piotr, Zhehuai Chen +17 · 1 citation
Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Sound (cs.SD) #Speech and dialogue systems #Topic Modeling #electronic engineering #information engineering
- SPGISpeech 2.0: Transcribed multi-speaker financial audio for speaker-tagged transcription
2025/08/07 by Grossman, Raymond, Park, Taejin, Dhawan, Kunal +6 · 1 citation
#Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
- Large Language Model Based Generative Error Correction: A Challenge and Baselines for Speech Recognition, Speaker Tagging, and Emotion Recognition
2024/09/15 by Chao-Han Huck Yang, Taejin Park, Yang, Chao-Han Huck +39 · 1 citation
Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #electronic engineering #information engineering
- Longer is (Not Necessarily) Stronger: Punctuated Long-Sequence Training for Enhanced Speech Recognition and Translation
2024/09/09 by Nithin Rao Koluguri, Koluguri, Nithin Rao, Travis Bartley +11 · 1 citation
Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Speech Recognition and Synthesis #electronic engineering #information engineering
- VoiceTextBlender: Augmenting Large Language Models with Speech Capabilities via Single-Stage Joint Speech-Text Supervised Fine-Tuning
2024/10/23 by Yifan Peng, Peng, Yifan, Krishna C. Puvvada +17 · 1 citation
Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Speech Recognition and Synthesis #Topic Modeling #electronic engineering #information engineering
- Streaming Sortformer: Speaker Cache-Based Online Speaker Diarization with Arrival-Time Ordering
2025/07/24 by Ivan Medennikov, Tae‐Jin Park, Medennikov, Ivan +13 · 3 citations
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering