2025/11/01 by Zhaoyan Zhang · 1 voice
Computer Science · Medicine · Psychology · #Speech Recognition and Synthesis #Stuttering Research and Treatment #Voice and Speech Disorders
paper · pdf · doi:10.1121/10.0039842
openalex publication_date 2025/11/01 · openalex created_date 2025/11/11 · openalex updated_date 2026/08/01
Currently, diagnosis of voice disorders is often made when patients visit the clinic, by which time speakers already experience vocal difficulties. The goal of this study was to develop a voice inversion system that predicts how speakers modulate vocal physiology from the produced voice, toward early detection of unhealthy vocal behavior. Two neural networks, a Bayesian neural network and a deep ensemble of neural networks, were developed that predict changes in vocal physiological parameters and their confidence intervals. Comparison to human data showed that both networks were able to predict meaningful differences in vocal behavior across subjects, demonstrating their potential toward ambulatory monitoring of vocal behavior at the physiological level.