Gaur, Yashesh
- The Llama 3 Herd of Models
2024/07/31 by Grattafiori, Aaron, Dubey, Abhimanyu, Jauhri, Abhinav +556 · 3039 citations
#Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences
- On decoder-only architecture for speech-to-text and large language model integration
2023/07/08 by Jian Wu, Wu, Jian, Yashesh Gaur +19 · 26 citations
Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Topic Modeling #electronic engineering #information engineering
- Serialized Output Training for End-to-End Overlapped Speech Recognition
2020/03/28 by Kanda, Naoyuki, Gaur, Yashesh, Wang, Xiaofei +2 · 13 citations
#Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
- Joint Speaker Counting, Speech Recognition, and Speaker Identification for Overlapped Speech of Any Number of Speakers
2020/06/19 by Kanda, Naoyuki, Gaur, Yashesh, Wang, Xiaofei +4 · 5 citations
#Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
- Streaming Speaker-Attributed ASR with Token-Level Speaker Embeddings
2022/03/30 by Kanda, Naoyuki, Wu, Jian, Wu, Yu +7 · 5 citations
#Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
- End-to-End Speaker-Attributed ASR with Transformer
2021/04/05 by Kanda, Naoyuki, Ye, Guoli, Gaur, Yashesh +4 · 4 citations
#Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
- Streaming Multi-Talker ASR with Token-Level Serialized Output Training
2022/02/02 by Kanda, Naoyuki, Wu, Jian, Wu, Yu +7 · 4 citations
#Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
- On the Comparison of Popular End-to-End Models for Large Scale Speech Recognition
2020/05/28 by Jinyu Li, Li, Jinyu, Yu Wu +9 · 3 citations
Computer Science · #Speech Recognition and Synthesis #Music and Audio Processing #Natural Language Processing Techniques
- VioLA: Unified Codec Language Models for Speech Recognition, Synthesis, and Translation
2023/05/25 by Wang, Tianrui, Zhou, Long, Zhang, Ziqiang +6 · 4 citations
#Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
- Large-Scale Streaming End-to-End Speech Translation with Neural Transducers
2022/04/11 by Xue, Jian, Wang, Peidong, Li, Jinyu +2 · 3 citations
#Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #electronic engineering #information engineering
- Transcribe-to-Diarize: Neural Speaker Diarization for Unlimited Number of Speakers using End-to-End Speaker-Attributed ASR
2021/10/07 by Kanda, Naoyuki, Xiao, Xiong, Gaur, Yashesh +4 · 2 citations
#Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
- A Comparative Study of Modular and Joint Approaches for Speaker-Attributed ASR on Monaural Long-Form Audio
2021/07/06 by Kanda, Naoyuki, Xiao, Xiong, Wu, Jian +6 · 2 citations
#Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
- Frozen Large Language Models Can Perceive Paralinguistic Aspects of Speech
2024/10/02 by Wonjune Kang, Kang, Wonjune, Junteng Jia +19 · 4 citations
Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Topic Modeling #electronic engineering #information engineering
- Exploring Neural Transducers for End-to-End Speech Recognition
2017/07/24 by Battenberg, Eric, Chen, Jitong, Child, Rewon +8 · 1 citation
#Computation and Language (cs.CL) #FOS: Computer and information sciences #Neural and Evolutionary Computing (cs.NE)
- COSMIC: Data Efficient Instruction-tuning For Speech In-Context Learning
2023/11/03 by Pan, Jing, Wu, Jian, Gaur, Yashesh +4 · 2 citations
#Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #electronic engineering #information engineering
- Large-Scale Pre-Training of End-to-End Multi-Talker ASR for Meeting Transcription with Single Distant Microphone
2021/03/31 by Kanda, Naoyuki, Ye, Guoli, Wu, Yu +5 · 1 citation
#Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
- Continuous Streaming Multi-Talker ASR with Dual-path Transducers
2021/09/17 by Raj, Desh, Lu, Liang, Chen, Zhuo +2 · 1 citation
#Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
- Speech ReaLLM -- Real-time Streaming Speech Recognition with Multimodal LLMs by Teaching the Flow of Time
2024/06/13 by Frank Seide, Morrie Doulaty, Seide, Frank +9 · 2 citations
Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #Speech and dialogue systems #electronic engineering #information engineering
- Textless Streaming Speech-to-Speech Translation using Semantic Speech Tokens
2024/10/04 by Jinzheng Zhao, Zhao, Jinzheng, Niko Moritz +17 · 1 citation
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Electrical engineering #Natural Language Processing Techniques #Speech Recognition and Synthesis #Speech and dialogue systems #electronic engineering #information engineering