Baevski, Alexei
- The Llama 3 Herd of Models
2024/07/31 by Grattafiori, Aaron, Dubey, Abhimanyu, Jauhri, Abhinav +556 · 2825 citations
#Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences
- wav2vec 2.0: A Framework for Self-Supervised Learning of Speech Representations
2020/06/20 by Baevski, Alexei, Zhou, Henry, Mohamed, Abdelrahman +1 · 506 citations
#Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Sound (cs.SD) #electronic engineering #information engineering
- wav2vec: Unsupervised Pre-training for Speech Recognition
2019/04/11 by Steffen Schneider, Schneider, Steffen, Alexei Baevski +5 · 1 voice · 17 citations
#cs.CL
- XLS-R: Self-supervised Cross-lingual Speech Representation Learning at Scale
2021/11/17 by Arun Babu, Changhan Wang, Babu, Arun +23 · 77 citations
Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #electronic engineering #information engineering
- fairseq: A Fast, Extensible Toolkit for Sequence Modeling
2019/04/01 by Myle Ott, Sergey Edunov, Ott, Myle +13 · 67 citations
Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Multimodal Machine Learning Applications #Natural Language Processing Techniques #Topic Modeling
- data2vec: A General Framework for Self-supervised Learning in Speech, Vision and Language
2022/02/07 by Alexei Baevski, Baevski, Alexei, Wei-Ning Hsu +9 · 62 citations
Computer Science · #Domain Adaptation and Few-Shot Learning #Speech Recognition and Synthesis #Multimodal Machine Learning Applications
- Scaling Speech Technology to 1,000+ Languages
2023/05/22 by Vineel Pratap, Pratap, Vineel, Andros Tjandra +29 · 71 citations
Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #electronic engineering #information engineering
- Unsupervised Cross-lingual Representation Learning for Speech\n Recognition
2020/06/24 by Alexis Conneau, Alexei Baevski, Conneau, Alexis +7 · 46 citations
Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- Generative Spoken Language Modeling from Raw Audio
2021/02/01 by Lakhotia, Kushal, Kharitonov, Evgeny, Hsu, Wei-Ning +8 · 39 citations
#Computation and Language (cs.CL) #FOS: Computer and information sciences
- Masked Autoencoders that Listen
2022/07/13 by Po-Yao Huang, Xu Hu, Huang, Po-Yao +13 · 42 citations
Computer Science · #Music and Audio Processing #Speech and Audio Processing #Speech Recognition and Synthesis
- Pay Less Attention with Lightweight and Dynamic Convolutions
2019/01/29 by Felix Wu, Wu, Felix, Angela Fan +7 · 41 citations
Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Multimodal Machine Learning Applications #Natural Language Processing Techniques #Topic Modeling
- vq-wav2vec: Self-Supervised Learning of Discrete Speech Representations
2019/10/12 by Alexei Baevski, Baevski, Alexei, Steffen Schneider +3 · 15 citations
Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Music and Audio Processing #Natural Language Processing Techniques #Speech Recognition and Synthesis
- Adaptive Input Representations for Neural Language Modeling
2018/09/28 by Alexei Baevski, Baevski, Alexei, Michael Auli +1 · 9 citations
Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Natural Language Processing Techniques #Speech Recognition and Synthesis #Topic Modeling
- The Zero Resource Speech Benchmark 2021: Metrics and baselines for unsupervised spoken language modeling
2020/11/23 by Nguyen, Tu Anh, de Seyssel, Maureen, Rozé, Patricia +5 · 10 citations
#Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
- Simple and Effective Zero-shot Cross-lingual Phoneme Recognition
2021/09/23 by Qiantong Xu, Xu, Qiantong, Alexei Baevski +3 · 10 citations
Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and dialogue systems
- Efficient Self-supervised Learning with Contextualized Target Representations for Vision, Speech and Language
2022/12/14 by Baevski, Alexei, Arun Babu, Wei-Ning Hsu +4 · 12 citations
Biochemistry, Genetics and Molecular Biology · Computer Science · #Audio and Speech Processing (eess.AS) #Cancer-related molecular mechanisms research #Computation and Language (cs.CL) #Domain Adaptation and Few-Shot Learning #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Sound (cs.SD) #Speech Recognition and Synthesis #electronic engineering #information engineering
- Offline Visual Representation Learning for Embodied Navigation
2022/04/27 by Karmesh Yadav, Yadav, Karmesh, Ram Ramrakhya +13 · 10 citations
Computer Science · #Multimodal Machine Learning Applications #Domain Adaptation and Few-Shot Learning #Human Pose and Action Recognition
- Robust wav2vec 2.0: Analyzing Domain Shift in Self-Supervised Pre-Training
2021/04/02 by Hsu, Wei-Ning, Sriram, Anuroop, Baevski, Alexei +8 · 8 citations
#Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Sound (cs.SD) #electronic engineering #information engineering
- OVRL-V2: A simple state-of-art baseline for ImageNav and ObjectNav
2023/03/14 by Yadav, Karmesh, Majumdar, Arjun, Ramrakhya, Ram +5 · 10 citations
#Artificial Intelligence (cs.AI) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences
- Facebook FAIR's WMT19 News Translation Task Submission
2019/07/15 by Ng, Nathan, Yee, Kyra, Baevski, Alexei +3 · 5 citations
#Computation and Language (cs.CL) #FOS: Computer and information sciences
- Unsupervised Speech Recognition
2021/05/24 by Alexei Baevski, Baevski, Alexei, Wei-Ning Hsu +5 · 5 citations
Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- Reservoir Transformers
2020/12/30 by Sheng Shen, Alexei Baevski, Shen, Sheng +9 · 4 citations
Computer Science · #Neural Networks and Reservoir Computing #Neural Networks and Applications #Machine Learning and ELM
- Towards End-to-end Unsupervised Speech Recognition
2022/04/05 by Liu, Alexander H., Hsu, Wei-Ning, Auli, Michael +1 · 4 citations
#Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
- Effectiveness of self-supervised pre-training for speech recognition
2019/11/10 by Baevski, Alexei, Auli, Michael, Mohamed, Abdelrahman · 3 citations
#Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG)
- AV-data2vec: Self-supervised Learning of Audio-Visual Speech Representations with Contextualized Target Representations
2023/02/10 by Jiachen Lian, Alexei Baevski, Lian, Jiachen +5 · 4 citations
Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- Unified Speech-Text Pre-training for Speech Translation and Recognition
2022/04/11 by Yun Tang, Tang, Yun, Hongyu Gong +19 · 3 citations
Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Natural Language Processing Techniques #Speech Recognition and Synthesis #Topic Modeling
- Self-training and Pre-training are Complementary for Speech Recognition
2020/10/22 by Xu, Qiantong, Baevski, Alexei, Likhomanenko, Tatiana +5 · 2 citations
#Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Sound (cs.SD) #electronic engineering #information engineering
- Simple and Effective Unsupervised Speech Synthesis
2022/04/06 by Liu, Alexander H., Lai, Cheng-I Jeff, Hsu, Wei-Ning +3 · 1 citation
#Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
- Toward Joint Language Modeling for Speech Units and Text
2023/10/12 by Ju-Chieh Chou, Chou, Ju-Chieh, Chung-Ming Chien +13 · 2 citations
Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Topic Modeling #electronic engineering #information engineering
- Multilingual Speech Translation with Efficient Finetuning of Pretrained Models
2020/10/24 by Xian Li, Li, Xian, Changhan Wang +15 · 2 citations
Computer Science · #Natural Language Processing Techniques #Speech Recognition and Synthesis #Topic Modeling