vix.ing · top · new · best · stats · spec

Fadi Biadsy

  1. Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
    2024/03/08 by Gemini Robotics Team, Petko Georgiev, Gemini Team +2277 · 4 voices · 1010 citations
    Computer Science · #Semantic Web and Ontologies
  2. Direct speech-to-speech translation with a sequence-to-sequence model
    2019/04/12 by Jia Ye, Jia, Ye, Ron J. Weiss +11 · 17 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Topic Modeling #electronic engineering #information engineering
  3. Parrotron: An End-to-End Speech-to-Speech Conversion Model and its Applications to Hearing-Impaired Speech and Speech Separation
    2019/04/08 by Fadi Biadsy, Biadsy, Fadi, Ron J. Weiss +7 · 3 citations
    Computer Science · Psychology · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Phonetics and Phonology Research #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  4. Residual Adapters for Parameter-Efficient ASR Adaptation to Atypical and\n Accented Speech
    2021/09/14 by Katrin Tomanek, Vicky Zayats, Tomanek, Katrin +7 · 2 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #Speech and dialogue systems #Topic Modeling #electronic engineering #information engineering
  5. Extending Multilingual Speech Synthesis to 100+ Languages without Transcribed Data
    2024/02/29 by Takaaki Saeki, Saeki, Takaaki, Gary Wang +19 · 2 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and dialogue systems #electronic engineering #information engineering
  6. Zero-shot Cross-lingual Voice Transfer for TTS
    2024/09/20 by Fadi Biadsy, Youzheng Chen, Biadsy, Fadi +11 · 1 citation
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering