Seltzer, Michael L.
- The Llama 3 Herd of Models
2024/07/31 by Grattafiori, Aaron, Dubey, Abhimanyu, Jauhri, Abhinav +556 · 3162 citations
#Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences
- Contextualized Streaming End-to-End Speech Recognition with Trie-Based Deep Biasing and Shallow Fusion
2021/04/05 by Duc Le, Le, Duc, Mahaveer Jain +21 · 12 citations
Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Natural Language Processing Techniques #Speech Recognition and Synthesis #Topic Modeling #electronic engineering #information engineering
- Semantic Distance: A New Metric for ASR Performance Analysis Towards Spoken Language Understanding
2021/04/05 by Suyoun Kim, Abhinav Arora, Kim, Suyoun +11 · 10 citations
Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Natural Language Processing Techniques #Speech Recognition and Synthesis #Topic Modeling
- Deep Shallow Fusion for RNN-T Personalization
2020/11/16 by Le, Duc, Keren, Gil, Chan, Julian +3 · 5 citations
#Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #electronic engineering #information engineering
- Towards measuring fairness in speech recognition: Fair-Speech dataset
2024/08/22 by Irina-Elena Veliche, Veliche, Irina-Elena, Zhuangqun Huang +9 · 10 citations
Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computers and Society (cs.CY) #FOS: Computer and information sciences #FOS: Electrical engineering #Hate Speech and Cyberbullying Detection #Machine Learning (stat.ML) #Sound (cs.SD) #electronic engineering #information engineering
- Massively Multilingual ASR on 70 Languages: Tokenization, Architecture, and Generalization Capabilities
2022/11/10 by Andros Tjandra, Nayan Singhal, Tjandra, Andros +11 · 5 citations
Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Multimodal Machine Learning Applications #Natural Language Processing Techniques #Sound (cs.SD) #Speech and dialogue systems #electronic engineering #information engineering
- End-to-End Speech Recognition Contextualization with Large Language Models
2023/09/19 by Egor Lakomkin, Chunyang Wu, Lakomkin, Egor +9 · 6 citations
Computer Science · #Speech Recognition and Synthesis #Natural Language Processing Techniques #Topic Modeling
- G2G: TTS-Driven Pronunciation Learning for Graphemic Hybrid ASR
2019/10/22 by Le, Duc, Koehler, Thilo, Fuegen, Christian +1 · 2 citations
#Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Sound (cs.SD) #electronic engineering #information engineering
- Transformer-Transducer: End-to-End Speech Recognition with Self-Attention
2019/10/28 by Yeh, Ching-Feng, Mahadeokar, Jay, Kalgaonkar, Kaustubh +6 · 2 citations
#Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
- Evaluating User Perception of Speech Recognition System Quality with Semantic Distance Metric
2021/10/11 by Kim, Suyoun, Le, Duc, Zheng, Weiyi +6 · 2 citations
#Computation and Language (cs.CL) #FOS: Computer and information sciences
- Improved training for online end-to-end speech recognition systems
2017/11/06 by Kim, Suyoun, Seltzer, Michael L., Li, Jinyu +1 · 1 citation
#Computation and Language (cs.CL) #FOS: Computer and information sciences
- End-to-end contextual speech recognition using class language models and a token passing decoder
2018/12/05 by Chen, Zhehuai, Jain, Mahaveer, Wang, Yongqiang +2 · 1 citation
#68T10 #Computation and Language (cs.CL) #FOS: Computer and information sciences #I.2.7 #Machine Learning (cs.LG)
- RNN-T For Latency Controlled ASR With Improved Beam Search
2019/11/05 by Jain, Mahaveer, Schubert, Kjell, Mahadeokar, Jay +5 · 1 citation
#Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #electronic engineering #information engineering
- Improved Neural Language Model Fusion for Streaming Recurrent Neural Network Transducer
2020/10/26 by Kim, Suyoun, Shangguan, Yuan, Mahadeokar, Jay +4 · 1 citation
#Computation and Language (cs.CL) #FOS: Computer and information sciences
- Dissecting User-Perceived Latency of On-Device E2E Speech Recognition
2021/04/06 by Shangguan, Yuan, Prabhavalkar, Rohit, Su, Hang +8 · 1 citation
#Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
- Factorized Blank Thresholding for Improved Runtime Efficiency of Neural Transducers
2022/11/02 by Manh Duc Le, Le, Duc, Frank Seide +11 · 1 citation
Computer Science · Earth and Planetary Sciences · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Neural Networks and Applications #Sound (cs.SD) #Speech Recognition and Synthesis #Underwater Acoustics Research #electronic engineering #information engineering
- Improving Fast-slow Encoder based Transducer with Streaming Deliberation
2022/12/15 by Ke Li, Li, Ke, Jay Mahadeokar +13 · 1 citation
Computer Science · Engineering · #Audio and Speech Processing (eess.AS) #FOS: Electrical engineering #Speech Recognition and Synthesis #Speech and Audio Processing #Ultrasonics and Acoustic Wave Propagation #electronic engineering #information engineering