vix.ing · top · new · best · stats · spec

Arun Babu

  1. Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model
    2024/08/20 by Chunting Zhou, Lili Yu, Zhou, Chunting +17 · 4 voices · 115 citations
    Neuroscience · #Brain Tumor Detection and Classification
  2. XLS-R: Self-supervised Cross-lingual Speech Representation Learning at Scale
    2021/11/17 by Arun Babu, Changhan Wang, Babu, Arun +23 · 77 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #electronic engineering #information engineering
  3. data2vec: A General Framework for Self-supervised Learning in Speech, Vision and Language
    2022/02/07 by Alexei Baevski, Wei-Ning Hsu, Baevski, Alexei +9 · 61 citations
    Computer Science · #Domain Adaptation and Few-Shot Learning #Speech Recognition and Synthesis #Multimodal Machine Learning Applications
  4. Scaling Speech Technology to 1,000+ Languages
    2023/05/22 by Vineel Pratap, Andros Tjandra, Pratap, Vineel +29 · 71 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #electronic engineering #information engineering
  5. Efficient Self-supervised Learning with Contextualized Target Representations for Vision, Speech and Language
    2022/12/14 by Baevski, Alexei, Arun Babu, Babu, Arun +4 · 12 citations
    Biochemistry, Genetics and Molecular Biology · Computer Science · #Audio and Speech Processing (eess.AS) #Cancer-related molecular mechanisms research #Computation and Language (cs.CL) #Domain Adaptation and Few-Shot Learning #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Sound (cs.SD) #Speech Recognition and Synthesis #electronic engineering #information engineering
  6. Scaling Autoregressive Multi-Modal Models: Pretraining and Instruction Tuning
    2023/09/05 by Lili Yu, Bowen Shi, Yu, Lili +51 · 7 citations
    Computer Science · #Multimodal Machine Learning Applications #Topic Modeling #Natural Language Processing Techniques
  7. Toward Joint Language Modeling for Speech Units and Text
    2023/10/12 by Ju-Chieh Chou, Chou, Ju-Chieh, Chung-Ming Chien +13 · 2 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Topic Modeling #electronic engineering #information engineering