vix.ing · top · new · best · stats · spec

Wei-Ning Hsu

  1. data2vec: A General Framework for Self-supervised Learning in Speech, Vision and Language
    2022/02/07 by Alexei Baevski, Wei-Ning Hsu, Baevski, Alexei +9 · 61 citations
    Computer Science · #Domain Adaptation and Few-Shot Learning #Speech Recognition and Synthesis #Multimodal Machine Learning Applications
  2. Movie Gen: A Cast of Media Foundation Models
    2024/10/17 by Adam Polyak, Polyak, Adam, Amit Zohar +164 · 123 citations
    Economics, Econometrics and Finance · #Cinema and Media Studies
  3. Scaling Speech Technology to 1,000+ Languages
    2023/05/22 by Vineel Pratap, Pratap, Vineel, Andros Tjandra +29 · 69 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #electronic engineering #information engineering
  4. Voicebox: Text-Guided Multilingual Universal Speech Generation at Scale
    2023/06/23 by Matthew Le, Apoorv Vyas, Le, Matthew +19 · 66 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Topic Modeling #electronic engineering #information engineering
  5. Learning Audio-Visual Speech Representation by Masked Multimodal Cluster Prediction
    2022/01/05 by Bowen Shi, Shi, Bowen, Wei-Ning Hsu +5 · 35 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #FOS: Electrical engineering #Face recognition and analysis #Sound (cs.SD) #Speech and Audio Processing #Video Analysis and Summarization #electronic engineering #information engineering
  6. Speech Resynthesis from Discrete Disentangled Self-Supervised Representations
    2021/04/01 by Adam Polyak, Polyak, Adam, Yossi Adi +13 · 20 citations
    Computer Science · #Speech and Audio Processing #Speech Recognition and Synthesis #Music and Audio Processing
  7. EXPRESSO: A Benchmark and Analysis of Discrete Expressive Speech Resynthesis
    2023/08/10 by Tu Anh Nguyen, Nguyen, Tu Anh, Wei-Ning Hsu +23 · 24 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and dialogue systems #Topic Modeling #electronic engineering #information engineering
  8. Tango 2: Aligning Diffusion-based Text-to-Audio Generations through Direct Preference Optimization
    2024/04/15 by Navonil Majumder, Chia-Yu Hung, Majumder, Navonil +9 · 29 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Music Technology and Sound Studies #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #electronic engineering #information engineering
  9. Meta Audiobox Aesthetics: Unified Automatic Quality Assessment for Speech, Music, and Sound
    2025/02/07 by Andros Tjandra, Yi-Chiao Wu, Tjandra, Andros +23 · 49 citations
    Computer Science · Psychology · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Multisensory perception and integration #Music Technology and Sound Studies #Music and Audio Processing #Sound (cs.SD) #electronic engineering #information engineering
  10. Audiobox: Unified Audio Generation with Natural Language Prompts
    2023/12/25 by Apoorv Vyas, Vyas, Apoorv, Bowen Shi +45 · 19 citations
    Computer Science · #Music and Audio Processing #Music Technology and Sound Studies #Speech Recognition and Synthesis
  11. Scaling Laws for Generative Mixed-Modal Language Models
    2023/01/10 by Armen Aghajanyan, Lili Yu, Aghajanyan, Armen +17 · 14 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Natural Language Processing Techniques #Speech Recognition and Synthesis #Topic Modeling
  12. Generative Pre-training for Speech with Flow Matching
    2023/10/25 by Alexander H. Liu, Liu, Alexander H., Matt Le +9 · 17 citations
    Computer Science · #Speech Recognition and Synthesis #Topic Modeling #Speech and dialogue systems
  13. Efficient Self-supervised Learning with Contextualized Target Representations for Vision, Speech and Language
    2022/12/14 by Arun Babu, Baevski, Alexei, Babu, Arun +4 · 12 citations
    Biochemistry, Genetics and Molecular Biology · Computer Science · #Audio and Speech Processing (eess.AS) #Cancer-related molecular mechanisms research #Computation and Language (cs.CL) #Domain Adaptation and Few-Shot Learning #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Sound (cs.SD) #Speech Recognition and Synthesis #electronic engineering #information engineering
  14. Robust Self-Supervised Audio-Visual Speech Recognition
    2022/01/05 by Bowen Shi, Shi, Bowen, Wei-Ning Hsu +3 · 10 citations
    Computer Science · Engineering · #Advanced Adaptive Filtering Techniques #Audio and Speech Processing (eess.AS) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Music and Audio Processing #Sound (cs.SD) #Speech and Audio Processing #electronic engineering #information engineering
  15. Text-Free Prosody-Aware Generative Spoken Language Modeling
    2021/09/07 by Eugene Kharitonov, Kharitonov, Eugene, Ann Lee +19 · 8 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Topic Modeling #electronic engineering #information engineering
  16. An Unsupervised Autoregressive Model for Speech Representation Learning
    2019/04/05 by Yu-An Chung, Wei-Ning Hsu, Chung, Yu-An +5 · 6 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  17. DinoSR: Self-Distillation and Online Clustering for Self-supervised Speech Representation Learning
    2023/05/17 by Alexander H. Liu, Liu, Alexander H., Heng-Jui Chang +7 · 8 citations
    Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Music and Audio Processing #Speech Recognition and Synthesis #Speech and dialogue systems
  18. Textless Speech-to-Speech Translation on Real Data
    2021/12/15 by Ann Lee, Hongyu Gong, Lee, Ann +18 · 5 citations
    Computer Science · #Speech Recognition and Synthesis #Natural Language Processing Techniques #Topic Modeling
  19. Textless Speech Emotion Conversion using Discrete and Decomposed Representations
    2021/11/14 by Felix Kreuk, Adam Polyak, Kreuk, Felix +16 · 5 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  20. Lingvo: a Modular and Scalable Framework for Sequence-to-Sequence Modeling
    2019/02/21 by Jonathan Shen, Shen, Jonathan, Patrick Nguyen +179 · 1 voice · 1 citation
    Computer Science · Mathematics · #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML) #cs.LG #stat.ML
  21. Action2Sound: Ambient-Aware Generation of Action Sounds from Egocentric Videos
    2024/06/13 by Changan Chen, Chen, Changan, Puyuan Peng +10 · 9 citations
    Computer Science · Engineering · #Human Pose and Action Recognition #Human Motion and Animation #Video Analysis and Summarization
  22. Unsupervised Learning of Disentangled and Interpretable Representations from Sequential Data
    2017/09/22 by Wei-Ning Hsu, Yu Zhang, Hsu, Wei-Ning +3 · 4 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Topic Modeling #electronic engineering #information engineering
  23. MuAViC: A Multilingual Audio-Visual Corpus for Robust Speech Recognition and Robust Speech-to-Text Translation
    2023/03/01 by Mohamed Anwar, Bowen Shi, Anwar, Mohamed +9 · 5 citations
    Computer Science · #Speech and Audio Processing #Speech Recognition and Synthesis #Music and Audio Processing
  24. u-HuBERT: Unified Mixed-Modal Speech Pretraining And Zero-Shot Transfer to Unlabeled Modality
    2022/07/14 by Wei-Ning Hsu, Hsu, Wei-Ning, Bowen Shi +1 · 4 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #FOS: Electrical engineering #Image and Video Processing (eess.IV) #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  25. High Fidelity Text-Guided Music Editing via Single-Stage Flow Matching
    2024/07/04 by Gaël Le Lan, Lan, Gael Le, Bowen Shi +21 · 7 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music Technology and Sound Studies #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #electronic engineering #information engineering
  26. AV-data2vec: Self-supervised Learning of Audio-Visual Speech Representations with Contextualized Target Representations
    2023/02/10 by Jiachen Lian, Lian, Jiachen, Alexei Baevski +5 · 4 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  27. Speech-to-Speech Translation For A Real-world Unwritten Language
    2022/11/11 by Peng‐Jen Chen, Kevin Tran, Chen, Peng-Jen +29 · 3 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Topic Modeling #electronic engineering #information engineering
  28. Unified Speech-Text Pre-training for Speech Translation and Recognition
    2022/04/11 by Yun Tang, Hongyu Gong, Tang, Yun +19 · 3 citations
    Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Natural Language Processing Techniques #Speech Recognition and Synthesis #Topic Modeling
  29. MusicFlow: Cascaded Flow Matching for Text Guided Music Generation
    2024/10/27 by K R Prajwal, Prajwal, K R, Bowen Shi +19 · 6 citations
    Computer Science · #Music and Audio Processing #Music Technology and Sound Studies #Natural Language Processing Techniques
  30. Cocktail HuBERT: Generalized Self-Supervised Pre-training for Mixture and Single-Source Speech
    2023/03/20 by Maryam Fazel-Zarandi, Fazel-Zarandi, Maryam, Wei-Ning Hsu +1 · 3 citations
    Computer Science · #Speech Recognition and Synthesis #Speech and Audio Processing #Music and Audio Processing
  31. FlowDec: A flow-based full-band general audio codec with high perceptual quality
    2025/03/03 by Simon Welker, Welker, Simon, Matthew Le +11 · 9 citations
    Computer Science · Engineering · #Advanced Data Compression Techniques #Speech and Audio Processing #Advanced Adaptive Filtering Techniques
  32. fairseq S2: A Scalable and Integrable Speech Synthesis Toolkit
    2021/09/14 by Changhan Wang, Wei-Ning Hsu, Wang, Changhan +13 · 2 citations
    Computer Science · #Speech Recognition and Synthesis #Natural Language Processing Techniques #Topic Modeling
  33. Disentangling by Partitioning: A Representation Learning Framework for\n Multimodal Sensory Data
    2018/05/29 by Wei-Ning Hsu, Hsu, Wei-Ning, James Glass +1 · 2 citations
    Computer Science · #Speech and Audio Processing
  34. textless-lib: a Library for Textless Spoken Language Processing
    2022/02/15 by Eugene Kharitonov, Kharitonov, Eugene, Jade Copet +19 · 1 citation
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and dialogue systems #electronic engineering #information engineering