Chiu, Chung-Cheng
- Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
2024/03/08 by Gemini Robotics Team, Petko Georgiev, Gemini Team +2277 · 4 voices · 579 citations
Computer Science · #Semantic Web and Ontologies
- Conformer: Convolution-augmented Transformer for Speech Recognition
2020/05/16 by Anmol Gulati, James Qin, Gulati, Anmol +19 · 224 citations
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- Apple Intelligence Foundation Language Models
2024/07/29 by Tom Gunter, Zirui Wang, Gunter, Tom +306 · 2 voices · 2 citations
#cs.AI #cs.CL #cs.LG
- Gemini: A Family of Highly Capable Multimodal Models
2023/12/19 by Gemini Robotics Team, Rohan Anil, Gemini Team +2692 · 9 voices · 6 citations
Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Multimodal Machine Learning Applications #Natural Language Processing Techniques #cs.AI #cs.CL #cs.CV
- W2v-BERT: Combining Contrastive Learning and Masked Language Modeling for Self-Supervised Speech Pre-Training
2021/08/07 by Yu-An Chung, Chung, Yu-An, Yu Zhang +11 · 33 citations
Computer Science · #Speech Recognition and Synthesis #Natural Language Processing Techniques #Topic Modeling
- Self-supervised Learning with Random-projection Quantizer for Speech Recognition
2022/02/03 by Chung‐Cheng Chiu, James Qin, Chiu, Chung-Cheng +7 · 28 citations
Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- Google USM: Scaling Automatic Speech Recognition Beyond 100 Languages
2023/03/02 by Yu Zhang, Zhang, Yu, Wei Han +51 · 27 citations
Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Topic Modeling #electronic engineering #information engineering
- Monotonic Chunkwise Attention
2017/12/14 by Chung‐Cheng Chiu, Colin Raffel, Chiu, Chung-Cheng +1 · 13 citations
Computer Science · #Speech Recognition and Synthesis #Topic Modeling #Natural Language Processing Techniques
- Minimum Word Error Rate Training for Attention-based Sequence-to-Sequence Models
2017/12/05 by Prabhavalkar, Rohit, Sainath, Tara N., Wu, Yonghui +4 · 8 citations
#Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (stat.ML) #electronic engineering #information engineering
- Monotonic Infinite Lookback Attention for Simultaneous Machine Translation
2019/06/12 by Arivazhagan, Naveen, Cherry, Colin, Macherey, Wolfgang +5 · 8 citations
#Computation and Language (cs.CL) #FOS: Computer and information sciences
- SLM: Bridge the thin gap between speech and text foundation models
2023/09/30 by Mingqiu Wang, Wei Han, Wang, Mingqiu +33 · 11 citations
Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Topic Modeling #electronic engineering #information engineering
- Two-Pass End-to-End Speech Recognition
2019/08/29 by Sainath, Tara N., Pang, Ruoming, Rybach, David +9 · 6 citations
#Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
- SpecAugment on Large Scale Datasets
2019/12/11 by Daniel Park, Yu Zhang, Park, Daniel S. +13 · 7 citations
Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- A comparison of end-to-end models for long-form speech recognition
2019/11/06 by Chung‐Cheng Chiu, Wei Han, Chiu, Chung-Cheng +25 · 7 citations
Computer Science · #Speech Recognition and Synthesis #Topic Modeling #Speech and dialogue systems
- A Comparison of Techniques for Language Model Integration in\n Encoder-Decoder Speech Recognition
2018/07/27 by Shubham Toshniwal, Anjuli Kannan, Toshniwal, Shubham +9 · 4 citations
Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- Talking Turns: Benchmarking Audio Foundation Models on Turn-Taking Dynamics
2025/03/03 by Siddhant Arora, Zhiyun Lu, Arora, Siddhant +7 · 1 voice · 12 citations
Psychology · Computer Science · #Phonetics and Phonology Research #Music Technology and Sound Studies #Music and Audio Processing
- Speech Sentiment Analysis via Pre-trained Features from End-to-end ASR Models
2019/11/21 by Lu, Zhiyun, Cao, Liangliang, Zhang, Yu +2 · 4 citations
#Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #electronic engineering #information engineering
- Lingvo: a Modular and Scalable Framework for Sequence-to-Sequence Modeling
2019/02/21 by Jonathan Shen, Shen, Jonathan, Patrick Nguyen +179 · 1 voice · 1 citation
Computer Science · Mathematics · #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML) #cs.LG #stat.ML
- A Better and Faster End-to-End Model for Streaming ASR
2020/11/21 by Li, Bo, Gulati, Anmol, Yu, Jiahui +12 · 4 citations
#Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
- State-of-the-art Speech Recognition With Sequence-to-Sequence Models
2017/12/05 by Chiu, Chung-Cheng, Sainath, Tara N., Wu, Yonghui +11 · 3 citations
#Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (stat.ML) #Sound (cs.SD) #electronic engineering #information engineering
- ContextNet: Improving Convolutional Neural Networks for Automatic Speech Recognition with Global Context
2020/05/07 by Wei Han, Zhengdong Zhang, Han, Wei +15 · 3 citations
Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- Pushing the Limits of Semi-Supervised Learning for Automatic Speech Recognition
2020/10/20 by Zhang, Yu, Qin, James, Park, Daniel S. +5 · 2 citations
#Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Sound (cs.SD) #electronic engineering #information engineering
- FastEmit: Low-latency Streaming ASR with Sequence-level Emission Regularization
2020/10/21 by Yu, Jiahui, Chiu, Chung-Cheng, Li, Bo +8 · 2 citations
#Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Sound (cs.SD) #electronic engineering #information engineering
- Apple Intelligence Foundation Language Models: Tech Report 2025
2025/07/17 by Ethan Li, Li, Ethan, Anders Larsen +487 · 14 citations
Computer Science · #Advanced Malware Detection Techniques #Advanced Neural Network Applications #Artificial Intelligence (cs.AI) #Big Data and Digital Economy #FOS: Computer and information sciences #Machine Learning (cs.LG)
- Accented Speech Recognition: Benchmarking, Pre-training, and Diverse Data
2022/05/16 by Alëna Aksënova, Aksënova, Alëna, Zhehuai Chen +19 · 2 citations
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and dialogue systems #electronic engineering #information engineering
- Recognizing long-form speech using streaming end-to-end models
2019/10/24 by Narayanan, Arun, Prabhavalkar, Rohit, Chiu, Chung-Cheng +3 · 1 citation
#Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
- Cascaded encoders for unifying streaming and non-streaming ASR
2020/10/27 by Narayanan, Arun, Sainath, Tara N., Pang, Ruoming +5 · 1 citation
#Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
- Efficient Knowledge Distillation for RNN-Transducer Models
2020/11/11 by Panchapagesan, Sankaran, Park, Daniel S., Chiu, Chung-Cheng +3 · 1 citation
#Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
- Parallel Rescoring with Transformer for Streaming On-Device Speech Recognition
2020/08/30 by Li, Wei, Qin, James, Chiu, Chung-Cheng +2 · 1 citation
#Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #electronic engineering #information engineering
- Textless Direct Speech-to-Speech Translation with Discrete Speech Representation
2022/10/31 by Xinjian Li, Jia Ye, Li, Xinjian +3 · 1 citation
Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Topic Modeling #electronic engineering #information engineering