Yanzhang He
- Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
2025/07/07 by Gheorghe Comanici, Comanici, Gheorghe, Eric Bieber +6844 · 8 voices · 1398 citations
#cs.CL #cs.AI
- Gemma 4 Technical Report
2026/07/02 by Gemma Team, Sherif El Abd, Vaibhav Aggarwal +320 · 7 voices · 8 citations
#cs.CL #cs.AI
- Confidence Estimation for Attention-based Sequence-to-sequence Models for Speech Recognition
2020/10/22 by Qiujia Li, David Qiu, Li, Qiujia +13 · 6 citations
Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Music and Audio Processing #Natural Language Processing Techniques #Speech Recognition and Synthesis #electronic engineering #information engineering
- Lingvo: a Modular and Scalable Framework for Sequence-to-Sequence Modeling
2019/02/21 by Jonathan Shen, Patrick Nguyen, Shen, Jonathan +179 · 1 voice · 1 citation
Computer Science · Mathematics · #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML) #cs.LG #stat.ML
- Learning Word-Level Confidence For Subword End-to-End ASR
2021/03/11 by David Qiu, Qiu, David, Qiujia Li +21 · 3 citations
Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Speech Recognition and Synthesis #Speech and dialogue systems #Topic Modeling #electronic engineering #information engineering
- Personal VAD 2.0: Optimizing Personal Voice Activity Detection for On-Device Speech Recognition
2022/04/08 by Shaojin Ding, Ding, Shaojin, Rajeev Rikhye +13 · 3 citations
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- Towards Fast and Accurate Streaming End-to-End ASR
2020/04/24 by Bo Li, Li, Bo, Shuo-Yiin Chang +11 · 2 citations
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Electrical engineering #Music and Audio Processing #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- VoiceFilter-Lite: Streaming Targeted Voice Separation for On-Device Speech Recognition
2020/09/09 by Quan Wang, Ignacio López Moreno, Wang, Quan +19 · 2 citations
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Music and Audio Processing #Signal Processing (eess.SP) #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- Streaming Small-Footprint Keyword Spotting using Sequence-to-Sequence Models
2017/10/26 by Yanzhang He, He, Yanzhang, Rohit Prabhavalkar +9 · 1 citation
Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Music and Audio Processing #Natural Language Processing Techniques #Speech Recognition and Synthesis
- Less Is More: Improved RNN-T Decoding Using Limited Label Context and\n Path Merging
2020/12/12 by Rohit Prabhavalkar, Yanzhang He, Prabhavalkar, Rohit +11 · 1 citation
Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- Multi-Task Learning for End-to-End ASR Word and Utterance Confidence with Deletion Prediction
2021/04/26 by David Qiu, Qiu, David, Yanzhang He +9 · 1 citation
Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Topic Modeling #electronic engineering #information engineering
- Personalized Keyphrase Detection using Speaker and Environment\n Information
2021/04/28 by Rajeev Rikhye, Rikhye, Rajeev, Quan Wang +14 · 1 citation
Computer Science · #Advanced Text Analysis Techniques #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Sound (cs.SD) #Text and Document Classification Technologies #electronic engineering #information engineering
- Large-scale ASR Domain Adaptation using Self- and Semi-supervised Learning
2021/10/01 by Dongseong Hwang, Hwang, Dongseong, Ananya Misra +17 · 1 citation
Biochemistry, Genetics and Molecular Biology · Computer Science · #Audio and Speech Processing (eess.AS) #Cancer-related molecular mechanisms research #Computation and Language (cs.CL) #Domain Adaptation and Few-Shot Learning #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Multimodal Machine Learning Applications #Sound (cs.SD) #electronic engineering #information engineering
- Turn-Taking Prediction for Natural Conversational Speech
2022/08/29 by Shuo-Yiin Chang, Chang, Shuo-yiin, Bo Li +11 · 1 citation
Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and dialogue systems #electronic engineering #information engineering