Sainath, Tara
- Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
2025/07/07 by Gheorghe Comanici, Eric Bieber, Comanici, Gheorghe +6844 · 8 voices · 1350 citations
#cs.CL #cs.AI
- Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
2024/03/08 by Gemini Robotics Team, Petko Georgiev, Gemini Team +2277 · 4 voices · 538 citations
Computer Science · #Semantic Web and Ontologies
- AudioPaLM: A Large Language Model That Can Speak and Listen
2023/06/22 by Paul K. Rubenstein, Chulayuth Asawaroengchai, Rubenstein, Paul K. +57 · 1 voice · 46 citations
#cs.CL #cs.AI #cs.SD #eess.AS #stat.ML
- Google USM: Scaling Automatic Speech Recognition Beyond 100 Languages
2023/03/02 by Yu Zhang, Zhang, Yu, Wei Han +51 · 26 citations
Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Topic Modeling #electronic engineering #information engineering
- Gemini: A Family of Highly Capable Multimodal Models
2023/12/19 by Gemini Team, Rohan Anil, Sebastian Borgeaud +2682 · 9 voices
Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #cs.AI #cs.CL #cs.CV
- A comparison of end-to-end models for long-form speech recognition
2019/11/06 by Chung‐Cheng Chiu, Wei Han, Chiu, Chung-Cheng +25 · 7 citations
Computer Science · #Speech Recognition and Synthesis #Topic Modeling #Speech and dialogue systems
- Lingvo: a Modular and Scalable Framework for Sequence-to-Sequence Modeling
2019/02/21 by Jonathan Shen, Shen, Jonathan, Patrick Nguyen +179 · 1 voice · 1 citation
Computer Science · Mathematics · #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML) #cs.LG #stat.ML
- Bytes are All You Need: End-to-End Multilingual Speech Recognition and Synthesis with Bytes
2018/11/22 by Li, Bo, Zhang, Yu, Sainath, Tara +2 · 2 citations
#Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Sound (cs.SD) #electronic engineering #information engineering
- Contextual Biasing with the Knuth-Morris-Pratt Matching Algorithm
2023/09/29 by Weiran Wang, Zelin Wu, Wang, Weiran +23 · 4 citations
Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Speech Recognition and Synthesis #Topic Modeling #electronic engineering #information engineering
- Deliberation of Streaming RNN-Transducer by Non-autoregressive Decoding
2021/12/01 by Weiran Wang, Ke Hu, Wang, Weiran +3 · 2 citations
Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Music and Audio Processing #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- Improving Speech Recognition for African American English With Audio Classification
2023/09/16 by Garg, Shefali, Huo, Zhouyuan, Sim, Khe Chai +11 · 2 citations
#Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Sound (cs.SD) #electronic engineering #information engineering