Wang, Gary
- Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
2024/03/08 by Gemini Robotics Team, Petko Georgiev, Gemini Team +2277 · 4 voices · 1056 citations
Computer Science · #Semantic Web and Ontologies
- Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
2025/07/07 by Gheorghe Comanici, Eric Bieber, Comanici, Gheorghe +6844 · 8 voices · 1459 citations
Computer Science · #cs.CL #cs.AI
- Google USM: Scaling Automatic Speech Recognition Beyond 100 Languages
2023/03/02 by Yu Zhang, Zhang, Yu, Wei Han +55 · 1 voice · 52 citations
Computer Science · Engineering · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Topic Modeling #cs.CL #cs.SD #eess.AS #electronic engineering #information engineering
- Accented Speech Recognition: Benchmarking, Pre-training, and Diverse Data
2022/05/16 by Alëna Aksënova, Aksënova, Alëna, Zhehuai Chen +19 · 3 citations
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and dialogue systems #electronic engineering #information engineering
- Extending Multilingual Speech Synthesis to 100+ Languages without Transcribed Data
2024/02/29 by Takaaki Saeki, Gary Wang, Saeki, Takaaki +19 · 2 citations
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and dialogue systems #electronic engineering #information engineering
- G-Augment: Searching for the Meta-Structure of Data Augmentation Policies for ASR
2022/10/19 by Wang, Gary, Cubuk, Ekin D., Rosenberg, Andrew +6 · 1 citation
#Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Sound (cs.SD) #electronic engineering #information engineering
- Virtuoso: Massive Multilingual Speech-Text Joint Semi-Supervised Learning for Text-To-Speech
2022/10/27 by Takaaki Saeki, Saeki, Takaaki, Heiga Zen +15 · 1 citation
Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and dialogue systems #electronic engineering #information engineering
- High-precision Voice Search Query Correction via Retrievable Speech-text Embedings
2024/01/08 by Christopher Li, Gary Wang, Li, Christopher +21 · 1 citation
Computer Science · #Speech Recognition and Synthesis #Speech and Audio Processing #Music and Audio Processing
- Zero-shot Cross-lingual Voice Transfer for TTS
2024/09/20 by Fadi Biadsy, Youzheng Chen, Biadsy, Fadi +11 · 1 citation
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering