Zhehuai Chen
- Google USM: Scaling Automatic Speech Recognition Beyond 100 Languages
2023/03/02 by Yu Zhang, Wei Han, Zhang, Yu +51 · 26 citations
Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Topic Modeling #electronic engineering #information engineering
- SALM: Speech-augmented Language Model with In-context Learning for Speech Recognition and Translation
2023/10/13 by Zhehuai Chen, Chen, Zhehuai, He Huang +15 · 13 citations
Computer Science · #68T10 #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Human-Computer Interaction (cs.HC) #I.2.7 #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and dialogue systems #electronic engineering #information engineering
- DeSTA2: Developing Instruction-Following Speech Language Model Without Speech Instruction-Tuning Data
2024/09/30 by Ke-Han Lu, Zhehuai Chen, Lu, Ke-Han +13 · 11 citations
Computer Science · #Speech and dialogue systems #Speech Recognition and Synthesis
- DeSTA: Enhancing Speech Language Models through Descriptive Speech-Text Alignment
2024/06/27 by Ke-Han Lu, Lu, Ke-Han, Zhehuai Chen +11 · 8 citations
Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Speech Recognition and Synthesis #Topic Modeling #electronic engineering #information engineering
- Chain-of-Thought Prompting for Speech Translation
2024/09/17 by Ke Hu, Zhehuai Chen, Hu, Ke +13 · 6 citations
Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Natural Language Processing Techniques
- Audio Large Language Models Can Be Descriptive Speech Quality Evaluators
2025/01/27 by Chen Chen, Chen, Chen, Yuchen Hu +13 · 8 citations
Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- SALM-Duplex: Efficient and Direct Duplex Modeling for Speech-to-Speech Language Model
2025/05/21 by Ke Hu, Hu, Ke, Ehsan Hosseini-Asl +17 · 11 citations
Computer Science · #Speech Recognition and Synthesis #Speech and dialogue systems #Speech and Audio Processing
- BESTOW: Efficient and Streamable Speech Language Model with the Best of Two Worlds in GPT and T5
2024/06/28 by Zhehuai Chen, He Huang, Chen, Zhehuai +13 · 4 citations
Computer Science · #68T10 #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Human-Computer Interaction (cs.HC) #I.2.7 #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Topic Modeling #electronic engineering #information engineering
- Accented Speech Recognition: Benchmarking, Pre-training, and Diverse Data
2022/05/16 by Alëna Aksënova, Aksënova, Alëna, Zhehuai Chen +19 · 2 citations
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and dialogue systems #electronic engineering #information engineering
- Accelerating RNN-T Training and Inference Using CTC guidance
2022/10/29 by Yongqiang Wang, Zhehuai Chen, Wang, Yongqiang +9 · 2 citations
Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Topic Modeling #electronic engineering #information engineering
- Modular End-to-end Automatic Speech Recognition Framework for Acoustic-to-word Model
2020/07/31 by Qi Liu, Liu, Qi, Zhehuai Chen +9 · 1 citation
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- JOIST: A Joint Speech and Text Streaming Model For ASR
2022/10/13 by Tara N. Sainath, Sainath, Tara N., Rohit Prabhavalkar +15 · 1 citation
Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Topic Modeling #electronic engineering #information engineering
- High-precision Voice Search Query Correction via Retrievable Speech-text Embedings
2024/01/08 by Christopher Li, Li, Christopher, Gary Wang +21 · 1 citation
Computer Science · #Speech Recognition and Synthesis #Speech and Audio Processing #Music and Audio Processing
- NeKo: Cross-Modality Post-Recognition Error Correction with Tasks-Guided Mixture-of-Experts Language Model
2024/11/08 by Yen‐Ting Lin, Lin, Yen-Ting, Zhehuai Chen +23 · 2 citations
Computer Science · #Topic Modeling #Natural Language Processing Techniques
- EMMeTT: Efficient Multimodal Machine Translation Training
2024/09/20 by Piotr Żelasko, Żelasko, Piotr, Zhehuai Chen +17 · 1 citation
Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Sound (cs.SD) #Speech and dialogue systems #Topic Modeling #electronic engineering #information engineering
- Large Language Model Based Generative Error Correction: A Challenge and Baselines for Speech Recognition, Speaker Tagging, and Emotion Recognition
2024/09/15 by Chao-Han Huck Yang, Yang, Chao-Han Huck, Taejin Park +39 · 1 citation
Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #electronic engineering #information engineering
- VoiceTextBlender: Augmenting Large Language Models with Speech Capabilities via Single-Stage Joint Speech-Text Supervised Fine-Tuning
2024/10/23 by Yifan Peng, Krishna C. Puvvada, Peng, Yifan +17 · 1 citation
Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Speech Recognition and Synthesis #Topic Modeling #electronic engineering #information engineering
- Just A Rather Very Intelligent Spoken Agent
2026/07/18 by Chen Chen, Zhehuai Chen
#cs.AI
- Voice Memory for Agentic Speech Recognition
2026/07/29 by Chao-Han Huck Yang, Zih-Ching Chen, Piotr Zelasko +3
Computer Science · Engineering · #cs.AI #cs.CL #cs.SD #eess.AS