vix.ing · top · new · best · stats · spec

Ramabhadran, Bhuvana

  1. Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
    2025/07/07 by Gheorghe Comanici, Eric Bieber, Comanici, Gheorghe +6844 · 8 voices · 1398 citations
    #cs.CL #cs.AI
  2. Google USM: Scaling Automatic Speech Recognition Beyond 100 Languages
    2023/03/02 by Yu Zhang, Wei Han, Zhang, Yu +55 · 1 voice · 38 citations
    Computer Science · Engineering · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Topic Modeling #cs.CL #cs.SD #eess.AS #electronic engineering #information engineering
  3. Speech Recognition with Augmented Synthesized Speech
    2019/09/25 by Rosenberg, Andrew, Zhang, Yu, Ramabhadran, Bhuvana +4 · 6 citations
    #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
  4. Learning to Speak Fluently in a Foreign Language: Multilingual Speech Synthesis and Cross-Language Voice Cloning
    2019/07/09 by Yu Zhang, Zhang, Yu, Ron J. Weiss +15 · 6 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  5. MAESTRO: Matched Speech Text Representations through Modality Matching
    2022/04/07 by Chen, Zhehuai, Zhang, Yu, Rosenberg, Andrew +4 · 5 citations
    #68T10 #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #I.2.7 #Sound (cs.SD) #electronic engineering #information engineering
  6. Large-Scale Multilingual Speech Recognition with a Streaming End-to-End Model
    2019/09/11 by Kannan, Anjuli, Datta, Arindrima, Sainath, Tara N. +6 · 3 citations
    #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Sound (cs.SD) #electronic engineering #information engineering
  7. English Conversational Telephone Speech Recognition by Humans and\n Machines
    2017/03/06 by George Saon, Saon, George, Gakuto Kurata +21 · 2 citations
    Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Natural Language Processing Techniques #Speech Recognition and Synthesis #Speech and Audio Processing
  8. Discrete Audio Tokens: More Than a Survey!
    2025/06/12 by Pooneh Mousavi, Mousavi, Pooneh, Gallil Maimon +39 · 13 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #Digital Media Forensic Detection #FOS: Computer and information sciences #FOS: Electrical engineering #Music Technology and Sound Studies #Music and Audio Processing #Sound (cs.SD) #electronic engineering #information engineering
  9. Accented Speech Recognition: Benchmarking, Pre-training, and Diverse Data
    2022/05/16 by Alëna Aksënova, Aksënova, Alëna, Zhehuai Chen +19 · 2 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and dialogue systems #electronic engineering #information engineering
  10. Improvements to deep convolutional neural networks for LVCSR
    2013/09/05 by Sainath, Tara N., Kingsbury, Brian, Mohamed, Abdel-rahman +6 · 1 citation
    #65K05 #90C15 #90C90 #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Mathematics #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Neural and Evolutionary Computing (cs.NE) #Optimization and Control (math.OC)
  11. Invariant Representations for Noisy Speech Recognition
    2016/11/27 by Serdyuk, Dmitriy, Audhkhasi, Kartik, Brakel, Philémon +3 · 1 citation
    #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Sound (cs.SD)
  12. Joint Modeling of Accents and Acoustics for Multi-Accent Speech Recognition
    2018/02/07 by Yang, Xuesong, Audhkhasi, Kartik, Rosenberg, Andrew +3 · 1 citation
    #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
  13. Building competitive direct acoustics-to-word models for English\n conversational speech recognition
    2017/12/08 by Kartik Audhkhasi, Audhkhasi, Kartik, Brian Kingsbury +7 · 1 citation
    Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (stat.ML) #Music and Audio Processing #Neural and Evolutionary Computing (cs.NE) #Speech Recognition and Synthesis #Speech and Audio Processing
  14. Direct Acoustics-to-Word Models for English Conversational Speech\n Recognition
    2017/03/22 by Kartik Audhkhasi, Bhuvana Ramabhadran, Audhkhasi, Kartik +7 · 2 citations
    Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (stat.ML) #Music and Audio Processing #Natural Language Processing Techniques #Neural and Evolutionary Computing (cs.NE) #Speech Recognition and Synthesis
  15. STAB: Speech Tokenizer Assessment Benchmark
    2024/09/04 by Shikhar Vashishth, Vashishth, Shikhar, Harman Preet Singh +15 · 2 citations
    Computer Science · #Speech Recognition and Synthesis
  16. Analysis of Self-Attention Head Diversity for Conformer-based Automatic Speech Recognition
    2022/09/13 by Audhkhasi, Kartik, Huang, Yinghui, Ramabhadran, Bhuvana +1 · 1 citation
    #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
  17. Extending Multilingual Speech Synthesis to 100+ Languages without Transcribed Data
    2024/02/29 by Takaaki Saeki, Saeki, Takaaki, Gary Wang +19 · 1 citation
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and dialogue systems #electronic engineering #information engineering
  18. Zero-shot Cross-lingual Voice Transfer for TTS
    2024/09/20 by Fadi Biadsy, Youzheng Chen, Biadsy, Fadi +11 · 1 citation
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  19. Schema Augmentation for Zero-Shot Domain Adaptation in Dialogue State Tracking
    2024/10/31 by Christopher Richardson, Richardson, Christopher, Roshan Sharma +9 · 1 citation
    Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Context-Aware Activity Recognition Systems #FOS: Computer and information sciences