vix.ing · top · new · best · stats · spec

Zhehuai Chen

  1. Google USM: Scaling Automatic Speech Recognition Beyond 100 Languages
    2023/03/02 by Yu Zhang, Wei Han, Zhang, Yu +51 · 26 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Topic Modeling #electronic engineering #information engineering
  2. SALM: Speech-augmented Language Model with In-context Learning for Speech Recognition and Translation
    2023/10/13 by Zhehuai Chen, Chen, Zhehuai, He Huang +15 · 13 citations
    Computer Science · #68T10 #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Human-Computer Interaction (cs.HC) #I.2.7 #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and dialogue systems #electronic engineering #information engineering
  3. DeSTA2: Developing Instruction-Following Speech Language Model Without Speech Instruction-Tuning Data
    2024/09/30 by Ke-Han Lu, Zhehuai Chen, Lu, Ke-Han +13 · 11 citations
    Computer Science · #Speech and dialogue systems #Speech Recognition and Synthesis
  4. DeSTA: Enhancing Speech Language Models through Descriptive Speech-Text Alignment
    2024/06/27 by Ke-Han Lu, Lu, Ke-Han, Zhehuai Chen +11 · 8 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Speech Recognition and Synthesis #Topic Modeling #electronic engineering #information engineering
  5. Chain-of-Thought Prompting for Speech Translation
    2024/09/17 by Ke Hu, Zhehuai Chen, Hu, Ke +13 · 6 citations
    Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Natural Language Processing Techniques
  6. Audio Large Language Models Can Be Descriptive Speech Quality Evaluators
    2025/01/27 by Chen Chen, Chen, Chen, Yuchen Hu +13 · 8 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  7. SALM-Duplex: Efficient and Direct Duplex Modeling for Speech-to-Speech Language Model
    2025/05/21 by Ke Hu, Hu, Ke, Ehsan Hosseini-Asl +17 · 11 citations
    Computer Science · #Speech Recognition and Synthesis #Speech and dialogue systems #Speech and Audio Processing
  8. BESTOW: Efficient and Streamable Speech Language Model with the Best of Two Worlds in GPT and T5
    2024/06/28 by Zhehuai Chen, He Huang, Chen, Zhehuai +13 · 4 citations
    Computer Science · #68T10 #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Human-Computer Interaction (cs.HC) #I.2.7 #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Topic Modeling #electronic engineering #information engineering
  9. Accented Speech Recognition: Benchmarking, Pre-training, and Diverse Data
    2022/05/16 by Alëna Aksënova, Aksënova, Alëna, Zhehuai Chen +19 · 2 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and dialogue systems #electronic engineering #information engineering
  10. Accelerating RNN-T Training and Inference Using CTC guidance
    2022/10/29 by Yongqiang Wang, Zhehuai Chen, Wang, Yongqiang +9 · 2 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Topic Modeling #electronic engineering #information engineering
  11. Modular End-to-end Automatic Speech Recognition Framework for Acoustic-to-word Model
    2020/07/31 by Qi Liu, Liu, Qi, Zhehuai Chen +9 · 1 citation
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  12. JOIST: A Joint Speech and Text Streaming Model For ASR
    2022/10/13 by Tara N. Sainath, Sainath, Tara N., Rohit Prabhavalkar +15 · 1 citation
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Topic Modeling #electronic engineering #information engineering
  13. High-precision Voice Search Query Correction via Retrievable Speech-text Embedings
    2024/01/08 by Christopher Li, Li, Christopher, Gary Wang +21 · 1 citation
    Computer Science · #Speech Recognition and Synthesis #Speech and Audio Processing #Music and Audio Processing
  14. NeKo: Cross-Modality Post-Recognition Error Correction with Tasks-Guided Mixture-of-Experts Language Model
    2024/11/08 by Yen‐Ting Lin, Lin, Yen-Ting, Zhehuai Chen +23 · 2 citations
    Computer Science · #Topic Modeling #Natural Language Processing Techniques
  15. EMMeTT: Efficient Multimodal Machine Translation Training
    2024/09/20 by Piotr Żelasko, Żelasko, Piotr, Zhehuai Chen +17 · 1 citation
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Sound (cs.SD) #Speech and dialogue systems #Topic Modeling #electronic engineering #information engineering
  16. Large Language Model Based Generative Error Correction: A Challenge and Baselines for Speech Recognition, Speaker Tagging, and Emotion Recognition
    2024/09/15 by Chao-Han Huck Yang, Yang, Chao-Han Huck, Taejin Park +39 · 1 citation
    Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #electronic engineering #information engineering
  17. VoiceTextBlender: Augmenting Large Language Models with Speech Capabilities via Single-Stage Joint Speech-Text Supervised Fine-Tuning
    2024/10/23 by Yifan Peng, Krishna C. Puvvada, Peng, Yifan +17 · 1 citation
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Speech Recognition and Synthesis #Topic Modeling #electronic engineering #information engineering
  18. Just A Rather Very Intelligent Spoken Agent
    2026/07/18 by Chen Chen, Zhehuai Chen
    #cs.AI
  19. Voice Memory for Agentic Speech Recognition
    2026/07/29 by Chao-Han Huck Yang, Zih-Ching Chen, Piotr Zelasko +3
    Computer Science · Engineering · #cs.AI #cs.CL #cs.SD #eess.AS