vix.ing · top · new · best · stats · spec

Mahadeokar, Jay

  1. The Llama 3 Herd of Models
    2024/07/31 by Grattafiori, Aaron, Dubey, Abhimanyu, Jauhri, Abhinav +556 · 2816 citations
    #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences
  2. Voicebox: Text-Guided Multilingual Universal Speech Generation at Scale
    2023/06/23 by Matthew Le, Le, Matthew, Apoorv Vyas +19 · 67 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Topic Modeling #electronic engineering #information engineering
  3. Prompting Large Language Models with Speech Recognition Abilities
    2023/07/21 by Yassir Fathullah, Fathullah, Yassir, Chunyang Wu +21 · 23 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Natural Language Processing Techniques #Speech Recognition and Synthesis #Topic Modeling #electronic engineering #information engineering
  4. Contextualized Streaming End-to-End Speech Recognition with Trie-Based Deep Biasing and Shallow Fusion
    2021/04/05 by Duc Le, Le, Duc, Mahaveer Jain +21 · 11 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Natural Language Processing Techniques #Speech Recognition and Synthesis #Topic Modeling #electronic engineering #information engineering
  5. Contextual RNN-T For Open Domain ASR
    2020/06/04 by Jain, Mahaveer, Keren, Gil, Mahadeokar, Jay +3 · 7 citations
    #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Sound (cs.SD) #electronic engineering #information engineering
  6. TorchAudio: Building Blocks for Audio and Speech Processing
    2021/10/28 by Yao-Yuan Yang, Moto Hira, Yang, Yao-Yuan +43 · 8 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #electronic engineering #information engineering
  7. AudioChatLlama: Towards General-Purpose Speech Abilities for LLMs
    2023/11/12 by Yassir Fathullah, Chunyang Wu, Fathullah, Yassir +16 · 11 citations
    Computer Science · #Speech Recognition and Synthesis #Topic Modeling #Speech and dialogue systems
  8. Deep Shallow Fusion for RNN-T Personalization
    2020/11/16 by Le, Duc, Keren, Gil, Chan, Julian +3 · 5 citations
    #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #electronic engineering #information engineering
  9. Transformer-Transducer: End-to-End Speech Recognition with Self-Attention
    2019/10/28 by Yeh, Ching-Feng, Mahadeokar, Jay, Kalgaonkar, Kaustubh +6 · 2 citations
    #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
  10. Frozen Large Language Models Can Perceive Paralinguistic Aspects of Speech
    2024/10/02 by Wonjune Kang, Kang, Wonjune, Junteng Jia +19 · 4 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Topic Modeling #electronic engineering #information engineering
  11. Efficient Streaming LLM for Speech Recognition
    2024/10/02 by Junteng Jia, Jia, Junteng, Gil Keren +15 · 3 citations
    Computer Science · Engineering · #Advanced Algorithms and Applications #Advanced Data Compression Techniques #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
  12. RNN-T For Latency Controlled ASR With Improved Beam Search
    2019/11/05 by Jain, Mahaveer, Schubert, Kjell, Mahadeokar, Jay +5 · 1 citation
    #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #electronic engineering #information engineering
  13. Improved Neural Language Model Fusion for Streaming Recurrent Neural Network Transducer
    2020/10/26 by Kim, Suyoun, Shangguan, Yuan, Mahadeokar, Jay +4 · 1 citation
    #Computation and Language (cs.CL) #FOS: Computer and information sciences
  14. Dissecting User-Perceived Latency of On-Device E2E Speech Recognition
    2021/04/06 by Shangguan, Yuan, Prabhavalkar, Rohit, Su, Hang +8 · 1 citation
    #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
  15. Federated Domain Adaptation for ASR with Full Self-Supervision
    2022/03/30 by Jia, Junteng, Mahadeokar, Jay, Zheng, Weiyi +3 · 1 citation
    #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
  16. Streaming Transformer Transducer Based Speech Recognition Using Non-Causal Convolution
    2021/10/07 by Shi, Yangyang, Wu, Chunyang, Wang, Dilin +9 · 1 citation
    #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #electronic engineering #information engineering
  17. An Investigation of Monotonic Transducers for Large-Scale Automatic Speech Recognition
    2022/04/19 by Moritz, Niko, Seide, Frank, Le, Duc +2 · 1 citation
    #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
  18. Dynamic Speech Endpoint Detection with Regression Targets
    2022/10/25 by Liang, Dawei, Su, Hang, Singh, Tarun +5 · 1 citation
    #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
  19. Improving Fast-slow Encoder based Transducer with Streaming Deliberation
    2022/12/15 by Li, Ke, Mahadeokar, Jay, Guo, Jinxi +5 · 1 citation
    #Audio and Speech Processing (eess.AS) #FOS: Electrical engineering #electronic engineering #information engineering
  20. TODM: Train Once Deploy Many Efficient Supernet-Based RNN-T Compression For On-device ASR Models
    2023/09/05 by Shangguan, Yuan, Yang, Haichuan, Li, Danni +11 · 1 citation
    #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Sound (cs.SD) #electronic engineering #information engineering
  21. Effective internal language model training and fusion for factorized transducer model
    2024/04/02 by Guo, Jinxi, Moritz, Niko, Ma, Yingyi +6 · 1 citation
    #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #electronic engineering #information engineering
  22. CJST: CTC Compressor based Joint Speech and Text Training for Decoder-Only ASR
    2024/11/12 by Zhou, Wei, Jia, Junteng, Sari, Leda +2 · 1 citation
    #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Sound (cs.SD) #electronic engineering #information engineering
  23. M-BEST-RQ: A Multi-Channel Speech Foundation Model for Smart Glasses
    2024/09/17 by Yang, Yufeng, Raj, Desh, Lin, Ju +8 · 1 citation
    #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering