vix.ing · top · new · best · stats · spec

Lee, Ann

  1. VoxPopuli: A Large-Scale Multilingual Speech Corpus for Representation Learning, Semi-Supervised Learning and Interpretation
    2021/01/02 by Changhan Wang, Morgane Rivière, Wang, Changhan +15 · 65 citations
    Computer Science · #Natural Language Processing Techniques
  2. Movie Gen: A Cast of Media Foundation Models
    2024/10/17 by Adam Polyak, Amit Zohar, Polyak, Adam +164 · 125 citations
    Economics, Econometrics and Finance · #Cinema and Media Studies
  3. Seamless: Multilingual Expressive and Streaming Speech Translation
    2023/12/08 by Seamless Communication, Loïc Barrault, Communication, Seamless +127 · 46 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Sound (cs.SD) #Speech and dialogue systems #Topic Modeling #electronic engineering #information engineering
  4. Meta Audiobox Aesthetics: Unified Automatic Quality Assessment for Speech, Music, and Sound
    2025/02/07 by Andros Tjandra, Tjandra, Andros, Yi-Chiao Wu +23 · 50 citations
    Computer Science · Psychology · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Multisensory perception and integration #Music Technology and Sound Studies #Music and Audio Processing #Sound (cs.SD) #electronic engineering #information engineering
  5. SeamlessM4T: Massively Multilingual & Multimodal Machine Translation
    2023/08/22 by Communication, Seamless, Barrault, Loïc, Chung, Yu-An +65 · 20 citations
    #Computation and Language (cs.CL) #FOS: Computer and information sciences #I.2.7
  6. Robust wav2vec 2.0: Analyzing Domain Shift in Self-Supervised Pre-Training
    2021/04/02 by Hsu, Wei-Ning, Sriram, Anuroop, Baevski, Alexei +8 · 8 citations
    #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Sound (cs.SD) #electronic engineering #information engineering
  7. Direct speech-to-speech translation with discrete units
    2021/07/12 by Lee, Ann, Chen, Peng-Jen, Wang, Changhan +9 · 8 citations
    #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #electronic engineering #information engineering
  8. Text-Free Prosody-Aware Generative Spoken Language Modeling
    2021/09/07 by Eugene Kharitonov, Kharitonov, Eugene, Ann Lee +19 · 8 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Topic Modeling #electronic engineering #information engineering
  9. UnitY: Two-pass Direct Speech-to-speech Translation with Discrete Units
    2022/12/15 by Hirofumi Inaguma, Sravya Popuri, Inaguma, Hirofumi +17 · 7 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  10. Sequence-to-Sequence Speech Recognition with Time-Depth Separable Convolutions
    2019/04/04 by Hannun, Awni, Lee, Ann, Xu, Qiantong +1 · 4 citations
    #Computation and Language (cs.CL) #FOS: Computer and information sciences
  11. Textless Speech-to-Speech Translation on Real Data
    2021/12/15 by Ann Lee, Lee, Ann, Hongyu Gong +18 · 5 citations
    Computer Science · #Speech Recognition and Synthesis #Natural Language Processing Techniques #Topic Modeling
  12. Flashlight: Enabling Innovation in Tools for Machine Learning
    2022/01/29 by Jacob Kahn, Vineel Pratap, Kahn, Jacob +25 · 4 citations
    Computer Science · Decision Sciences · #Artificial Intelligence (cs.AI) #Data Stream Mining Techniques #Distributed #Explainable Artificial Intelligence (XAI) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Parallel #Scientific Computing and Data Management #and Cluster Computing (cs.DC)
  13. Augmentation Invariant Discrete Representation for Generative Spoken Language Modeling
    2022/09/30 by Gat, Itai, Kreuk, Felix, Nguyen, Tu Anh +5 · 4 citations
    #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #electronic engineering #information engineering
  14. Speech-to-Speech Translation For A Real-world Unwritten Language
    2022/11/11 by Peng‐Jen Chen, Kevin Tran, Chen, Peng-Jen +29 · 3 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Topic Modeling #electronic engineering #information engineering
  15. fairseq S2: A Scalable and Integrable Speech Synthesis Toolkit
    2021/09/14 by Changhan Wang, Wei-Ning Hsu, Wang, Changhan +13 · 2 citations
    Computer Science · #Speech Recognition and Synthesis #Natural Language Processing Techniques #Topic Modeling
  16. Few-shot Sequence Learning with Transformers
    2020/12/17 by Logeswaran, Lajanugen, Lee, Ann, Ott, Myle +3 · 1 citation
    #FOS: Computer and information sciences #Machine Learning (cs.LG)
  17. Enhanced Direct Speech-to-Speech Translation Using Self-supervised Pre-training and Data Augmentation
    2022/04/06 by Popuri, Sravya, Chen, Peng-Jen, Wang, Changhan +5 · 1 citation
    #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
  18. Bridging Speech and Textual Pre-trained Models with Unsupervised ASR
    2022/11/06 by Jiatong Shi, Chan-Jan Hsu, Shi, Jiatong +13 · 1 citation
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Topic Modeling #electronic engineering #information engineering
  19. SpeechMatrix: A Large-Scale Mined Corpus of Multilingual Speech-to-Speech Translations
    2022/11/08 by Duquenne, Paul-Ambroise, Gong, Hongyu, Dong, Ning +7 · 1 citation
    #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
  20. textless-lib: a Library for Textless Spoken Language Processing
    2022/02/15 by Eugene Kharitonov, Jade Copet, Kharitonov, Eugene +19 · 1 citation
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and dialogue systems #electronic engineering #information engineering
  21. SAM Audio: Segment Anything in Audio
    2025/12/19 by Shi, Bowen, Tjandra, Andros, Hoffman, John +11 · 3 citations
    #Audio and Speech Processing (eess.AS) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #FOS: Electrical engineering #electronic engineering #information engineering
  22. Textless Acoustic Model with Self-Supervised Distillation for Noise-Robust Expressive Speech-to-Speech Translation
    2024/06/04 by Min-Jae Hwang, Hwang, Min-Jae, Ilia Kulikov +9 · 1 citation
    Computer Science · #Speech Recognition and Synthesis #Natural Language Processing Techniques #Topic Modeling
  23. Semi-Supervised Speech Recognition via Local Prior Matching
    2020/02/24 by Wei-Ning Hsu, Ann Lee, Hsu, Wei-Ning +5 · 1 citation
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Music and Audio Processing #Natural Language Processing Techniques #Speech Recognition and Synthesis #electronic engineering #information engineering