vix.ing · top · new · best · stats · spec

Automatic assessment of voice quality using machine learning

2026/02/25 by Yat Chun Au, Nan Yan, Manwa L. Ng · 1 voice
Computer Science · Medicine · #Phonocardiography and Auscultation Techniques #Speech Recognition and Synthesis #Voice and Speech Disorders

paper · doi:10.1080/14015439.2026.2628250

openalex publication_date 2026/02/25 · openalex created_date 2026/02/26 · openalex updated_date 2026/07/22

Abstract

Objectives This study aimed to develop and validate machine learning (ML) models for automated prediction of perceptual dysphonia severity, as indexed by the Grade (G) parameter of the GRBAS scale, using acoustic analyses of sustained vowels. The overarching goal was to enhance objectivity, reproducibility, and efficiency in clinical voice assessment.Methods A total of 524 sustained/a/samples were collected from three databases. The evaluations of all recordings by ten raters using the GRBAS scale, that achieved excellent interrater reliability (Krippendorff’s α = 0.96), were modelled. Forty-seven acoustic features spanning spectral, cepstral, perturbation, and noise-based indices were extracted using Parselmouth (Praat). Five ML classifiers—Decision Tree (DT), Random Forest (RF), Extreme Gradient Boosting (XGBoost), Light Gradient Boosting Machine (LightGBM), and Categorical Boosting (CatBoost)—were trained using 5-fold cross-validation (80/20 split) and evaluated by accuracy, F1-score, and quadratic weighted kappa (QWK).Results Gradient boosting algorithms outperformed traditional tree-based models. LightGBM achieved the highest QWK (0.945), followed by CatBoost (QWK = 0.941) and XGBoost (QWK = 0.935). Feature-importance analyses identified cepstral measures—particularly Smoothed Cepstral Peak Prominence (CPPS), Cepstral Spectral Index of Dysphonia (CSID), Acoustic Voice Quality Index (AVQI), Harmonics-to-Noise Ratio (HNR) as the most influential predictors of perceptual Grade (G), while jitter and shimmer parameters contributed minimally. Correlation analyses confirmed strong associations between Grade and AVQI (r = 0.854), HNR (r = −0.853), and cepstral indices (r = −0.835 to −0.832).Conclusions Gradient boosting methods, particularly LightGBM, produced near-expert agreement with perceptual ratings, supporting their potential as objective, interpretable tools for clinical dysphonia assessment.

Citations

Discussions

Related