Ramabhadran, Bhuvana
- Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
2025/07/07 by Gheorghe Comanici, Eric Bieber, Comanici, Gheorghe +6844 · 8 voices · 1398 citations
#cs.CL #cs.AI
- Google USM: Scaling Automatic Speech Recognition Beyond 100 Languages
2023/03/02 by Yu Zhang, Wei Han, Zhang, Yu +55 · 1 voice · 38 citations
Computer Science · Engineering · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Topic Modeling #cs.CL #cs.SD #eess.AS #electronic engineering #information engineering
- Speech Recognition with Augmented Synthesized Speech
2019/09/25 by Rosenberg, Andrew, Zhang, Yu, Ramabhadran, Bhuvana +4 · 6 citations
#Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
- Learning to Speak Fluently in a Foreign Language: Multilingual Speech Synthesis and Cross-Language Voice Cloning
2019/07/09 by Yu Zhang, Zhang, Yu, Ron J. Weiss +15 · 6 citations
Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- MAESTRO: Matched Speech Text Representations through Modality Matching
2022/04/07 by Chen, Zhehuai, Zhang, Yu, Rosenberg, Andrew +4 · 5 citations
#68T10 #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #I.2.7 #Sound (cs.SD) #electronic engineering #information engineering
- Large-Scale Multilingual Speech Recognition with a Streaming End-to-End Model
2019/09/11 by Kannan, Anjuli, Datta, Arindrima, Sainath, Tara N. +6 · 3 citations
#Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Sound (cs.SD) #electronic engineering #information engineering
- English Conversational Telephone Speech Recognition by Humans and\n Machines
2017/03/06 by George Saon, Saon, George, Gakuto Kurata +21 · 2 citations
Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Natural Language Processing Techniques #Speech Recognition and Synthesis #Speech and Audio Processing
- Discrete Audio Tokens: More Than a Survey!
2025/06/12 by Pooneh Mousavi, Mousavi, Pooneh, Gallil Maimon +39 · 13 citations
Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #Digital Media Forensic Detection #FOS: Computer and information sciences #FOS: Electrical engineering #Music Technology and Sound Studies #Music and Audio Processing #Sound (cs.SD) #electronic engineering #information engineering
- Accented Speech Recognition: Benchmarking, Pre-training, and Diverse Data
2022/05/16 by Alëna Aksënova, Aksënova, Alëna, Zhehuai Chen +19 · 2 citations
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and dialogue systems #electronic engineering #information engineering
- Improvements to deep convolutional neural networks for LVCSR
2013/09/05 by Sainath, Tara N., Kingsbury, Brian, Mohamed, Abdel-rahman +6 · 1 citation
#65K05 #90C15 #90C90 #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Mathematics #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Neural and Evolutionary Computing (cs.NE) #Optimization and Control (math.OC)
- Invariant Representations for Noisy Speech Recognition
2016/11/27 by Serdyuk, Dmitriy, Audhkhasi, Kartik, Brakel, Philémon +3 · 1 citation
#Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Sound (cs.SD)
- Joint Modeling of Accents and Acoustics for Multi-Accent Speech Recognition
2018/02/07 by Yang, Xuesong, Audhkhasi, Kartik, Rosenberg, Andrew +3 · 1 citation
#Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
- Building competitive direct acoustics-to-word models for English\n conversational speech recognition
2017/12/08 by Kartik Audhkhasi, Audhkhasi, Kartik, Brian Kingsbury +7 · 1 citation
Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (stat.ML) #Music and Audio Processing #Neural and Evolutionary Computing (cs.NE) #Speech Recognition and Synthesis #Speech and Audio Processing
- Direct Acoustics-to-Word Models for English Conversational Speech\n Recognition
2017/03/22 by Kartik Audhkhasi, Bhuvana Ramabhadran, Audhkhasi, Kartik +7 · 2 citations
Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (stat.ML) #Music and Audio Processing #Natural Language Processing Techniques #Neural and Evolutionary Computing (cs.NE) #Speech Recognition and Synthesis
- STAB: Speech Tokenizer Assessment Benchmark
2024/09/04 by Shikhar Vashishth, Vashishth, Shikhar, Harman Preet Singh +15 · 2 citations
Computer Science · #Speech Recognition and Synthesis
- Analysis of Self-Attention Head Diversity for Conformer-based Automatic Speech Recognition
2022/09/13 by Audhkhasi, Kartik, Huang, Yinghui, Ramabhadran, Bhuvana +1 · 1 citation
#Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
- Extending Multilingual Speech Synthesis to 100+ Languages without Transcribed Data
2024/02/29 by Takaaki Saeki, Saeki, Takaaki, Gary Wang +19 · 1 citation
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and dialogue systems #electronic engineering #information engineering
- Zero-shot Cross-lingual Voice Transfer for TTS
2024/09/20 by Fadi Biadsy, Youzheng Chen, Biadsy, Fadi +11 · 1 citation
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- Schema Augmentation for Zero-Shot Domain Adaptation in Dialogue State Tracking
2024/10/31 by Christopher Richardson, Richardson, Christopher, Roshan Sharma +9 · 1 citation
Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Context-Aware Activity Recognition Systems #FOS: Computer and information sciences