Thomas Drugman
- BASE TTS: Lessons from building a billion-parameter Text-to-Speech model on 100K hours of data
2024/02/12 by Mateusz Łajszczak, Łajszczak, Mateusz, Guillermo Cámbara +36 · 2 voices · 22 citations
Computer Science · Engineering · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Natural Language Processing Techniques #Speech Recognition and Synthesis #Topic Modeling #cs.CL #cs.LG #eess.AS #electronic engineering #information engineering
- Objective Study of Sensor Relevance for Automatic Cough Detection
2019/12/30 by Thomas Drugman, Drugman, Thomas, Jérôme Urbain +11 · 2 citations
Medicine · Health Professions · #Respiratory and Cough-Related Research #Voice and Speech Disorders #Infant Health and Development
- Weakly-supervised word-level pronunciation error detection in non-native\n English speech
2021/06/07 by Daniel Korzekwa, Jaime Lorenzo-Trueba, Korzekwa, Daniel +7 · 2 citations
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Natural Language Processing Techniques #Speech Recognition and Synthesis #Speech and dialogue systems #electronic engineering #information engineering
- Controllable Emphasis with zero data for text-to-speech
2023/07/13 by Arnaud Joly, Joly, Arnaud, Marco Nicolis +25 · 2 citations
Computer Science · Psychology · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Phonetics and Phonology Research #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- Interpretable Deep Learning Model for the Detection and Reconstruction\n of Dysarthric Speech
2019/07/10 by Daniel Korzekwa, Roberto Barra-Chicote, Korzekwa, Daniel +7 · 1 citation
Computer Science · Medicine · Psychology · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Phonetics and Phonology Research #Sound (cs.SD) #Speech Recognition and Synthesis #Voice and Speech Disorders #electronic engineering #information engineering
- A Comparative Study of Glottal Source Estimation Techniques
2019/12/28 by Thomas Drugman, Drugman, Thomas, Barış Bozkurt +3 · 1 citation
Computer Science · Medicine · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #Voice and Speech Disorders #electronic engineering #information engineering
- Joint Robust Voicing Detection and Pitch Estimation Based on Residual Harmonics
2019/12/28 by Thomas Drugman, Abeer Alwan, Drugman, Thomas +1 · 1 citation
Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- Detection of Glottal Closure Instants from Speech Signals: a\n Quantitative Review
2019/12/28 by Thomas Drugman, Mark R. Thomas, Drugman, Thomas +8 · 1 citation
Computer Science · Psychology · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Phonetics and Phonology Research #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- Prosodic Representation Learning and Contextual Sampling for Neural\n Text-to-Speech
2020/11/04 by Sri Karlapati, Ammar N. Abbas, Karlapati, Sri +11 · 1 citation
Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Topic Modeling #electronic engineering #information engineering
- Voicy: Zero-Shot Non-Parallel Voice Conversion in Noisy Reverberant\n Environments
2021/06/16 by Alejandro Mottini, Mottini, Alejandro, Jaime Lorenzo-Trueba +5 · 1 citation
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering