vix.ing · top · new · best · stats · spec

Baevski, Alexei

  1. The Llama 3 Herd of Models
    2024/07/31 by Grattafiori, Aaron, Dubey, Abhimanyu, Jauhri, Abhinav +556 · 2825 citations
    #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences
  2. wav2vec 2.0: A Framework for Self-Supervised Learning of Speech Representations
    2020/06/20 by Baevski, Alexei, Zhou, Henry, Mohamed, Abdelrahman +1 · 506 citations
    #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Sound (cs.SD) #electronic engineering #information engineering
  3. wav2vec: Unsupervised Pre-training for Speech Recognition
    2019/04/11 by Steffen Schneider, Schneider, Steffen, Alexei Baevski +5 · 1 voice · 17 citations
    #cs.CL
  4. XLS-R: Self-supervised Cross-lingual Speech Representation Learning at Scale
    2021/11/17 by Arun Babu, Changhan Wang, Babu, Arun +23 · 77 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #electronic engineering #information engineering
  5. fairseq: A Fast, Extensible Toolkit for Sequence Modeling
    2019/04/01 by Myle Ott, Sergey Edunov, Ott, Myle +13 · 67 citations
    Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Multimodal Machine Learning Applications #Natural Language Processing Techniques #Topic Modeling
  6. data2vec: A General Framework for Self-supervised Learning in Speech, Vision and Language
    2022/02/07 by Alexei Baevski, Baevski, Alexei, Wei-Ning Hsu +9 · 62 citations
    Computer Science · #Domain Adaptation and Few-Shot Learning #Speech Recognition and Synthesis #Multimodal Machine Learning Applications
  7. Scaling Speech Technology to 1,000+ Languages
    2023/05/22 by Vineel Pratap, Pratap, Vineel, Andros Tjandra +29 · 71 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #electronic engineering #information engineering
  8. Unsupervised Cross-lingual Representation Learning for Speech\n Recognition
    2020/06/24 by Alexis Conneau, Alexei Baevski, Conneau, Alexis +7 · 46 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  9. Generative Spoken Language Modeling from Raw Audio
    2021/02/01 by Lakhotia, Kushal, Kharitonov, Evgeny, Hsu, Wei-Ning +8 · 39 citations
    #Computation and Language (cs.CL) #FOS: Computer and information sciences
  10. Masked Autoencoders that Listen
    2022/07/13 by Po-Yao Huang, Xu Hu, Huang, Po-Yao +13 · 42 citations
    Computer Science · #Music and Audio Processing #Speech and Audio Processing #Speech Recognition and Synthesis
  11. Pay Less Attention with Lightweight and Dynamic Convolutions
    2019/01/29 by Felix Wu, Wu, Felix, Angela Fan +7 · 41 citations
    Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Multimodal Machine Learning Applications #Natural Language Processing Techniques #Topic Modeling
  12. vq-wav2vec: Self-Supervised Learning of Discrete Speech Representations
    2019/10/12 by Alexei Baevski, Baevski, Alexei, Steffen Schneider +3 · 15 citations
    Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Music and Audio Processing #Natural Language Processing Techniques #Speech Recognition and Synthesis
  13. Adaptive Input Representations for Neural Language Modeling
    2018/09/28 by Alexei Baevski, Baevski, Alexei, Michael Auli +1 · 9 citations
    Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Natural Language Processing Techniques #Speech Recognition and Synthesis #Topic Modeling
  14. The Zero Resource Speech Benchmark 2021: Metrics and baselines for unsupervised spoken language modeling
    2020/11/23 by Nguyen, Tu Anh, de Seyssel, Maureen, Rozé, Patricia +5 · 10 citations
    #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
  15. Simple and Effective Zero-shot Cross-lingual Phoneme Recognition
    2021/09/23 by Qiantong Xu, Xu, Qiantong, Alexei Baevski +3 · 10 citations
    Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and dialogue systems
  16. Efficient Self-supervised Learning with Contextualized Target Representations for Vision, Speech and Language
    2022/12/14 by Baevski, Alexei, Arun Babu, Wei-Ning Hsu +4 · 12 citations
    Biochemistry, Genetics and Molecular Biology · Computer Science · #Audio and Speech Processing (eess.AS) #Cancer-related molecular mechanisms research #Computation and Language (cs.CL) #Domain Adaptation and Few-Shot Learning #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Sound (cs.SD) #Speech Recognition and Synthesis #electronic engineering #information engineering
  17. Offline Visual Representation Learning for Embodied Navigation
    2022/04/27 by Karmesh Yadav, Yadav, Karmesh, Ram Ramrakhya +13 · 10 citations
    Computer Science · #Multimodal Machine Learning Applications #Domain Adaptation and Few-Shot Learning #Human Pose and Action Recognition
  18. Robust wav2vec 2.0: Analyzing Domain Shift in Self-Supervised Pre-Training
    2021/04/02 by Hsu, Wei-Ning, Sriram, Anuroop, Baevski, Alexei +8 · 8 citations
    #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Sound (cs.SD) #electronic engineering #information engineering
  19. OVRL-V2: A simple state-of-art baseline for ImageNav and ObjectNav
    2023/03/14 by Yadav, Karmesh, Majumdar, Arjun, Ramrakhya, Ram +5 · 10 citations
    #Artificial Intelligence (cs.AI) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences
  20. Facebook FAIR's WMT19 News Translation Task Submission
    2019/07/15 by Ng, Nathan, Yee, Kyra, Baevski, Alexei +3 · 5 citations
    #Computation and Language (cs.CL) #FOS: Computer and information sciences
  21. Unsupervised Speech Recognition
    2021/05/24 by Alexei Baevski, Baevski, Alexei, Wei-Ning Hsu +5 · 5 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  22. Reservoir Transformers
    2020/12/30 by Sheng Shen, Alexei Baevski, Shen, Sheng +9 · 4 citations
    Computer Science · #Neural Networks and Reservoir Computing #Neural Networks and Applications #Machine Learning and ELM
  23. Towards End-to-end Unsupervised Speech Recognition
    2022/04/05 by Liu, Alexander H., Hsu, Wei-Ning, Auli, Michael +1 · 4 citations
    #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
  24. Effectiveness of self-supervised pre-training for speech recognition
    2019/11/10 by Baevski, Alexei, Auli, Michael, Mohamed, Abdelrahman · 3 citations
    #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG)
  25. AV-data2vec: Self-supervised Learning of Audio-Visual Speech Representations with Contextualized Target Representations
    2023/02/10 by Jiachen Lian, Alexei Baevski, Lian, Jiachen +5 · 4 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  26. Unified Speech-Text Pre-training for Speech Translation and Recognition
    2022/04/11 by Yun Tang, Tang, Yun, Hongyu Gong +19 · 3 citations
    Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Natural Language Processing Techniques #Speech Recognition and Synthesis #Topic Modeling
  27. Self-training and Pre-training are Complementary for Speech Recognition
    2020/10/22 by Xu, Qiantong, Baevski, Alexei, Likhomanenko, Tatiana +5 · 2 citations
    #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Sound (cs.SD) #electronic engineering #information engineering
  28. Simple and Effective Unsupervised Speech Synthesis
    2022/04/06 by Liu, Alexander H., Lai, Cheng-I Jeff, Hsu, Wei-Ning +3 · 1 citation
    #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
  29. Toward Joint Language Modeling for Speech Units and Text
    2023/10/12 by Ju-Chieh Chou, Chou, Ju-Chieh, Chung-Ming Chien +13 · 2 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Topic Modeling #electronic engineering #information engineering
  30. Multilingual Speech Translation with Efficient Finetuning of Pretrained Models
    2020/10/24 by Xian Li, Li, Xian, Changhan Wang +15 · 2 citations
    Computer Science · #Natural Language Processing Techniques #Speech Recognition and Synthesis #Topic Modeling