vix.ing · top · new · best · stats · spec

Yanzhang He

  1. Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
    2025/07/07 by Gheorghe Comanici, Comanici, Gheorghe, Eric Bieber +6844 · 8 voices · 1398 citations
    #cs.CL #cs.AI
  2. Gemma 4 Technical Report
    2026/07/02 by Gemma Team, Sherif El Abd, Vaibhav Aggarwal +320 · 7 voices · 8 citations
    #cs.CL #cs.AI
  3. Confidence Estimation for Attention-based Sequence-to-sequence Models for Speech Recognition
    2020/10/22 by Qiujia Li, David Qiu, Li, Qiujia +13 · 6 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Music and Audio Processing #Natural Language Processing Techniques #Speech Recognition and Synthesis #electronic engineering #information engineering
  4. Lingvo: a Modular and Scalable Framework for Sequence-to-Sequence Modeling
    2019/02/21 by Jonathan Shen, Patrick Nguyen, Shen, Jonathan +179 · 1 voice · 1 citation
    Computer Science · Mathematics · #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML) #cs.LG #stat.ML
  5. Learning Word-Level Confidence For Subword End-to-End ASR
    2021/03/11 by David Qiu, Qiu, David, Qiujia Li +21 · 3 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Speech Recognition and Synthesis #Speech and dialogue systems #Topic Modeling #electronic engineering #information engineering
  6. Personal VAD 2.0: Optimizing Personal Voice Activity Detection for On-Device Speech Recognition
    2022/04/08 by Shaojin Ding, Ding, Shaojin, Rajeev Rikhye +13 · 3 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  7. Towards Fast and Accurate Streaming End-to-End ASR
    2020/04/24 by Bo Li, Li, Bo, Shuo-Yiin Chang +11 · 2 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Electrical engineering #Music and Audio Processing #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  8. VoiceFilter-Lite: Streaming Targeted Voice Separation for On-Device Speech Recognition
    2020/09/09 by Quan Wang, Ignacio López Moreno, Wang, Quan +19 · 2 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Music and Audio Processing #Signal Processing (eess.SP) #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  9. Streaming Small-Footprint Keyword Spotting using Sequence-to-Sequence Models
    2017/10/26 by Yanzhang He, He, Yanzhang, Rohit Prabhavalkar +9 · 1 citation
    Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Music and Audio Processing #Natural Language Processing Techniques #Speech Recognition and Synthesis
  10. Less Is More: Improved RNN-T Decoding Using Limited Label Context and\n Path Merging
    2020/12/12 by Rohit Prabhavalkar, Yanzhang He, Prabhavalkar, Rohit +11 · 1 citation
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  11. Multi-Task Learning for End-to-End ASR Word and Utterance Confidence with Deletion Prediction
    2021/04/26 by David Qiu, Qiu, David, Yanzhang He +9 · 1 citation
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Topic Modeling #electronic engineering #information engineering
  12. Personalized Keyphrase Detection using Speaker and Environment\n Information
    2021/04/28 by Rajeev Rikhye, Rikhye, Rajeev, Quan Wang +14 · 1 citation
    Computer Science · #Advanced Text Analysis Techniques #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Sound (cs.SD) #Text and Document Classification Technologies #electronic engineering #information engineering
  13. Large-scale ASR Domain Adaptation using Self- and Semi-supervised Learning
    2021/10/01 by Dongseong Hwang, Hwang, Dongseong, Ananya Misra +17 · 1 citation
    Biochemistry, Genetics and Molecular Biology · Computer Science · #Audio and Speech Processing (eess.AS) #Cancer-related molecular mechanisms research #Computation and Language (cs.CL) #Domain Adaptation and Few-Shot Learning #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Multimodal Machine Learning Applications #Sound (cs.SD) #electronic engineering #information engineering
  14. Turn-Taking Prediction for Natural Conversational Speech
    2022/08/29 by Shuo-Yiin Chang, Chang, Shuo-yiin, Bo Li +11 · 1 citation
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and dialogue systems #electronic engineering #information engineering