vix.ing · top · new · best · stats · spec

Sainath, Tara

  1. Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
    2025/07/07 by Gheorghe Comanici, Eric Bieber, Comanici, Gheorghe +6844 · 8 voices · 1350 citations
    #cs.CL #cs.AI
  2. Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
    2024/03/08 by Gemini Robotics Team, Petko Georgiev, Gemini Team +2277 · 4 voices · 538 citations
    Computer Science · #Semantic Web and Ontologies
  3. AudioPaLM: A Large Language Model That Can Speak and Listen
    2023/06/22 by Paul K. Rubenstein, Chulayuth Asawaroengchai, Rubenstein, Paul K. +57 · 1 voice · 46 citations
    #cs.CL #cs.AI #cs.SD #eess.AS #stat.ML
  4. Google USM: Scaling Automatic Speech Recognition Beyond 100 Languages
    2023/03/02 by Yu Zhang, Zhang, Yu, Wei Han +51 · 26 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Topic Modeling #electronic engineering #information engineering
  5. Gemini: A Family of Highly Capable Multimodal Models
    2023/12/19 by Gemini Team, Rohan Anil, Sebastian Borgeaud +2682 · 9 voices
    Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #cs.AI #cs.CL #cs.CV
  6. A comparison of end-to-end models for long-form speech recognition
    2019/11/06 by Chung‐Cheng Chiu, Wei Han, Chiu, Chung-Cheng +25 · 7 citations
    Computer Science · #Speech Recognition and Synthesis #Topic Modeling #Speech and dialogue systems
  7. Lingvo: a Modular and Scalable Framework for Sequence-to-Sequence Modeling
    2019/02/21 by Jonathan Shen, Shen, Jonathan, Patrick Nguyen +179 · 1 voice · 1 citation
    Computer Science · Mathematics · #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML) #cs.LG #stat.ML
  8. Bytes are All You Need: End-to-End Multilingual Speech Recognition and Synthesis with Bytes
    2018/11/22 by Li, Bo, Zhang, Yu, Sainath, Tara +2 · 2 citations
    #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Sound (cs.SD) #electronic engineering #information engineering
  9. Contextual Biasing with the Knuth-Morris-Pratt Matching Algorithm
    2023/09/29 by Weiran Wang, Zelin Wu, Wang, Weiran +23 · 4 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Speech Recognition and Synthesis #Topic Modeling #electronic engineering #information engineering
  10. Deliberation of Streaming RNN-Transducer by Non-autoregressive Decoding
    2021/12/01 by Weiran Wang, Ke Hu, Wang, Weiran +3 · 2 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Music and Audio Processing #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  11. Improving Speech Recognition for African American English With Audio Classification
    2023/09/16 by Garg, Shefali, Huo, Zhouyuan, Sim, Khe Chai +11 · 2 citations
    #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Sound (cs.SD) #electronic engineering #information engineering