Puvvada, Krishna C.
- SALM: Speech-augmented Language Model with In-context Learning for Speech Recognition and Translation
2023/10/13 by Zhehuai Chen, He Huang, Chen, Zhehuai +15 · 17 citations
Computer Science · #68T10 #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Human-Computer Interaction (cs.HC) #I.2.7 #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and dialogue systems #electronic engineering #information engineering
- Sortformer: Seamless Integration of Speaker Diarization and ASR by Bridging Timestamps and Tokens
2024/09/10 by Taejin Park, Ivan Medennikov, Park, Taejin +15 · 15 citations
Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and dialogue systems #electronic engineering #information engineering
- Discrete Audio Representation as an Alternative to Mel-Spectrograms for Speaker and Speech Recognition
2023/09/19 by Puvvada, Krishna C., Koluguri, Nithin Rao, Dhawan, Kunal +2 · 4 citations
#Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
- BESTOW: Efficient and Streamable Speech Language Model with the Best of Two Worlds in GPT and T5
2024/06/28 by Zhehuai Chen, He Huang, Chen, Zhehuai +13 · 5 citations
Computer Science · #68T10 #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Human-Computer Interaction (cs.HC) #I.2.7 #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Topic Modeling #electronic engineering #information engineering
- Training and Inference Efficiency of Encoder-Decoder Speech Models
2025/03/07 by Żelasko, Piotr, Dhawan, Kunal, Galvez, Daniel +7 · 6 citations
#Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #electronic engineering #information engineering
- NVIDIA Nemotron 3: Efficient and Open Intelligence
2025/12/24 by NVIDIA, :, Aaron Blakeman +453 · 5 citations
Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Explainable Artificial Intelligence (XAI) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Multimodal Machine Learning Applications #Topic Modeling
- SWAN-GPT: An Efficient and Scalable Approach for Long-Context Language Modeling
2025/04/11 by Krishna C. Puvvada, Puvvada, Krishna C., Faisal Ladhak +19 · 6 citations
Computer Science · #Topic Modeling #Natural Language Processing Techniques
- Few-shot acoustic event detection via meta-learning
2020/02/21 by Shi, Bowen, Sun, Ming, Puvvada, Krishna C. +3 · 1 citation
#Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Sound (cs.SD) #electronic engineering #information engineering
- VoiceTextBlender: Augmenting Large Language Models with Speech Capabilities via Single-Stage Joint Speech-Text Supervised Fine-Tuning
2024/10/23 by Yifan Peng, Peng, Yifan, Krishna C. Puvvada +17 · 3 citations
Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Speech Recognition and Synthesis #Topic Modeling #electronic engineering #information engineering
- Nemotron 3 Nano: Open, Efficient Mixture-of-Experts Hybrid Mamba-Transformer Model for Agentic Reasoning
2025/12/23 by NVIDIA, :, Aaron Blakeman +623 · 1 voice · 2 citations
#cs.CL #cs.AI #cs.LG
- Accidental Learners: Spoken Language Identification in Multilingual Self-Supervised Models
2022/11/09 by Travis Bartley, Fei Jia, Bartley, Travis M. +7 · 1 citation
Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #electronic engineering #information engineering
- NEST: Self-supervised Fast Conformer as All-purpose Seasoning to Speech Processing Tasks
2024/08/23 by Huang, He, Park, Taejin, Dhawan, Kunal +6 · 2 citations
#Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
- Multilingual Audio-Visual Speech Recognition with Hybrid CTC/RNN-T Fast Conformer
2024/03/14 by Maxime Burchi, Burchi, Maxime, Krishna C. Puvvada +7 · 1 citation
Computer Science · #Advanced Data Compression Techniques #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #FOS: Electrical engineering #Multimedia (cs.MM) #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- Resource-Efficient Adaptation of Speech Foundation Models for Multi-Speaker ASR
2024/09/02 by Wang, Weiqing, Dhawan, Kunal, Park, Taejin +6 · 1 citation
#Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering