vix.ing · top · new · best · stats · spec

Thomas Drugman

  1. BASE TTS: Lessons from building a billion-parameter Text-to-Speech model on 100K hours of data
    2024/02/12 by Mateusz Łajszczak, Łajszczak, Mateusz, Guillermo Cámbara +36 · 2 voices · 22 citations
    Computer Science · Engineering · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Natural Language Processing Techniques #Speech Recognition and Synthesis #Topic Modeling #cs.CL #cs.LG #eess.AS #electronic engineering #information engineering
  2. Objective Study of Sensor Relevance for Automatic Cough Detection
    2019/12/30 by Thomas Drugman, Drugman, Thomas, Jérôme Urbain +11 · 2 citations
    Medicine · Health Professions · #Respiratory and Cough-Related Research #Voice and Speech Disorders #Infant Health and Development
  3. Weakly-supervised word-level pronunciation error detection in non-native\n English speech
    2021/06/07 by Daniel Korzekwa, Jaime Lorenzo-Trueba, Korzekwa, Daniel +7 · 2 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Natural Language Processing Techniques #Speech Recognition and Synthesis #Speech and dialogue systems #electronic engineering #information engineering
  4. Controllable Emphasis with zero data for text-to-speech
    2023/07/13 by Arnaud Joly, Joly, Arnaud, Marco Nicolis +25 · 2 citations
    Computer Science · Psychology · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Phonetics and Phonology Research #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  5. Interpretable Deep Learning Model for the Detection and Reconstruction\n of Dysarthric Speech
    2019/07/10 by Daniel Korzekwa, Roberto Barra-Chicote, Korzekwa, Daniel +7 · 1 citation
    Computer Science · Medicine · Psychology · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Phonetics and Phonology Research #Sound (cs.SD) #Speech Recognition and Synthesis #Voice and Speech Disorders #electronic engineering #information engineering
  6. A Comparative Study of Glottal Source Estimation Techniques
    2019/12/28 by Thomas Drugman, Drugman, Thomas, Barış Bozkurt +3 · 1 citation
    Computer Science · Medicine · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #Voice and Speech Disorders #electronic engineering #information engineering
  7. Joint Robust Voicing Detection and Pitch Estimation Based on Residual Harmonics
    2019/12/28 by Thomas Drugman, Abeer Alwan, Drugman, Thomas +1 · 1 citation
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  8. Detection of Glottal Closure Instants from Speech Signals: a\n Quantitative Review
    2019/12/28 by Thomas Drugman, Mark R. Thomas, Drugman, Thomas +8 · 1 citation
    Computer Science · Psychology · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Phonetics and Phonology Research #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  9. Prosodic Representation Learning and Contextual Sampling for Neural\n Text-to-Speech
    2020/11/04 by Sri Karlapati, Ammar N. Abbas, Karlapati, Sri +11 · 1 citation
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Topic Modeling #electronic engineering #information engineering
  10. Voicy: Zero-Shot Non-Parallel Voice Conversion in Noisy Reverberant\n Environments
    2021/06/16 by Alejandro Mottini, Mottini, Alejandro, Jaime Lorenzo-Trueba +5 · 1 citation
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering