vix.ing · top · new · best · stats · spec

Wang, Gary

  1. Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
    2024/03/08 by Gemini Robotics Team, Petko Georgiev, Gemini Team +2277 · 4 voices · 1056 citations
    Computer Science · #Semantic Web and Ontologies
  2. Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
    2025/07/07 by Gheorghe Comanici, Eric Bieber, Comanici, Gheorghe +6844 · 8 voices · 1459 citations
    Computer Science · #cs.CL #cs.AI
  3. Google USM: Scaling Automatic Speech Recognition Beyond 100 Languages
    2023/03/02 by Yu Zhang, Zhang, Yu, Wei Han +55 · 1 voice · 52 citations
    Computer Science · Engineering · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Topic Modeling #cs.CL #cs.SD #eess.AS #electronic engineering #information engineering
  4. Accented Speech Recognition: Benchmarking, Pre-training, and Diverse Data
    2022/05/16 by Alëna Aksënova, Aksënova, Alëna, Zhehuai Chen +19 · 3 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and dialogue systems #electronic engineering #information engineering
  5. Extending Multilingual Speech Synthesis to 100+ Languages without Transcribed Data
    2024/02/29 by Takaaki Saeki, Gary Wang, Saeki, Takaaki +19 · 2 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and dialogue systems #electronic engineering #information engineering
  6. G-Augment: Searching for the Meta-Structure of Data Augmentation Policies for ASR
    2022/10/19 by Wang, Gary, Cubuk, Ekin D., Rosenberg, Andrew +6 · 1 citation
    #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Sound (cs.SD) #electronic engineering #information engineering
  7. Virtuoso: Massive Multilingual Speech-Text Joint Semi-Supervised Learning for Text-To-Speech
    2022/10/27 by Takaaki Saeki, Saeki, Takaaki, Heiga Zen +15 · 1 citation
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and dialogue systems #electronic engineering #information engineering
  8. High-precision Voice Search Query Correction via Retrievable Speech-text Embedings
    2024/01/08 by Christopher Li, Gary Wang, Li, Christopher +21 · 1 citation
    Computer Science · #Speech Recognition and Synthesis #Speech and Audio Processing #Music and Audio Processing
  9. Zero-shot Cross-lingual Voice Transfer for TTS
    2024/09/20 by Fadi Biadsy, Youzheng Chen, Biadsy, Fadi +11 · 1 citation
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering