vix.ing · top · new · best · stats · spec

Gaur, Yashesh

  1. The Llama 3 Herd of Models
    2024/07/31 by Grattafiori, Aaron, Dubey, Abhimanyu, Jauhri, Abhinav +556 · 3039 citations
    #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences
  2. On decoder-only architecture for speech-to-text and large language model integration
    2023/07/08 by Jian Wu, Wu, Jian, Yashesh Gaur +19 · 26 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Topic Modeling #electronic engineering #information engineering
  3. Serialized Output Training for End-to-End Overlapped Speech Recognition
    2020/03/28 by Kanda, Naoyuki, Gaur, Yashesh, Wang, Xiaofei +2 · 13 citations
    #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
  4. Joint Speaker Counting, Speech Recognition, and Speaker Identification for Overlapped Speech of Any Number of Speakers
    2020/06/19 by Kanda, Naoyuki, Gaur, Yashesh, Wang, Xiaofei +4 · 5 citations
    #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
  5. Streaming Speaker-Attributed ASR with Token-Level Speaker Embeddings
    2022/03/30 by Kanda, Naoyuki, Wu, Jian, Wu, Yu +7 · 5 citations
    #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
  6. End-to-End Speaker-Attributed ASR with Transformer
    2021/04/05 by Kanda, Naoyuki, Ye, Guoli, Gaur, Yashesh +4 · 4 citations
    #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
  7. Streaming Multi-Talker ASR with Token-Level Serialized Output Training
    2022/02/02 by Kanda, Naoyuki, Wu, Jian, Wu, Yu +7 · 4 citations
    #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
  8. On the Comparison of Popular End-to-End Models for Large Scale Speech Recognition
    2020/05/28 by Jinyu Li, Li, Jinyu, Yu Wu +9 · 3 citations
    Computer Science · #Speech Recognition and Synthesis #Music and Audio Processing #Natural Language Processing Techniques
  9. VioLA: Unified Codec Language Models for Speech Recognition, Synthesis, and Translation
    2023/05/25 by Wang, Tianrui, Zhou, Long, Zhang, Ziqiang +6 · 4 citations
    #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
  10. Large-Scale Streaming End-to-End Speech Translation with Neural Transducers
    2022/04/11 by Xue, Jian, Wang, Peidong, Li, Jinyu +2 · 3 citations
    #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #electronic engineering #information engineering
  11. Transcribe-to-Diarize: Neural Speaker Diarization for Unlimited Number of Speakers using End-to-End Speaker-Attributed ASR
    2021/10/07 by Kanda, Naoyuki, Xiao, Xiong, Gaur, Yashesh +4 · 2 citations
    #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
  12. A Comparative Study of Modular and Joint Approaches for Speaker-Attributed ASR on Monaural Long-Form Audio
    2021/07/06 by Kanda, Naoyuki, Xiao, Xiong, Wu, Jian +6 · 2 citations
    #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
  13. Frozen Large Language Models Can Perceive Paralinguistic Aspects of Speech
    2024/10/02 by Wonjune Kang, Kang, Wonjune, Junteng Jia +19 · 4 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Topic Modeling #electronic engineering #information engineering
  14. Exploring Neural Transducers for End-to-End Speech Recognition
    2017/07/24 by Battenberg, Eric, Chen, Jitong, Child, Rewon +8 · 1 citation
    #Computation and Language (cs.CL) #FOS: Computer and information sciences #Neural and Evolutionary Computing (cs.NE)
  15. COSMIC: Data Efficient Instruction-tuning For Speech In-Context Learning
    2023/11/03 by Pan, Jing, Wu, Jian, Gaur, Yashesh +4 · 2 citations
    #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #electronic engineering #information engineering
  16. Large-Scale Pre-Training of End-to-End Multi-Talker ASR for Meeting Transcription with Single Distant Microphone
    2021/03/31 by Kanda, Naoyuki, Ye, Guoli, Wu, Yu +5 · 1 citation
    #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
  17. Continuous Streaming Multi-Talker ASR with Dual-path Transducers
    2021/09/17 by Raj, Desh, Lu, Liang, Chen, Zhuo +2 · 1 citation
    #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
  18. Speech ReaLLM -- Real-time Streaming Speech Recognition with Multimodal LLMs by Teaching the Flow of Time
    2024/06/13 by Frank Seide, Morrie Doulaty, Seide, Frank +9 · 2 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #Speech and dialogue systems #electronic engineering #information engineering
  19. Textless Streaming Speech-to-Speech Translation using Semantic Speech Tokens
    2024/10/04 by Jinzheng Zhao, Zhao, Jinzheng, Niko Moritz +17 · 1 citation
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Electrical engineering #Natural Language Processing Techniques #Speech Recognition and Synthesis #Speech and dialogue systems #electronic engineering #information engineering