vix.ing · top · new · best · stats · spec

Speech based Depression Severity Level Classification Using a\n Multi-Stage Dilated CNN-LSTM Model

2021/04/09 by Nadee Seneviratne, Seneviratne, Nadee, Carol Espy-Wilson +1
Computer Science · Medicine · Psychology · #Audio and Speech Processing (eess.AS) #Emotion and Mood Recognition #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Speech Recognition and Synthesis #Voice and Speech Disorders #electronic engineering #information engineering

paper · pdf · doi:10.48550/arxiv.2104.04195

openalex publication_date 2021/04/09 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28

Abstract

Speech based depression classification has gained immense popularity over the\nrecent years. However, most of the classification studies have focused on\nbinary classification to distinguish depressed subjects from non-depressed\nsubjects. In this paper, we formulate the depression classification task as a\nseverity level classification problem to provide more granularity to the\nclassification outcomes. We use articulatory coordination features (ACFs)\ndeveloped to capture the changes of neuromotor coordination that happens as a\nresult of psychomotor slowing, a necessary feature of Major Depressive\nDisorder. The ACFs derived from the vocal tract variables (TVs) are used to\ntrain a dilated Convolutional Neural Network based depression classification\nmodel to obtain segment-level predictions. Then, we propose a Recurrent Neural\nNetwork based approach to obtain session-level predictions from segment-level\npredictions. We show that strengths of the segment-wise classifier are\namplified when a session-wise classifier is trained on embeddings obtained from\nit. The model trained on ACFs derived from TVs show relative improvement of\n27.47% in Unweighted Average Recall (UAR) at the session-level classification\ntask, compared to the ACFs derived from Mel Frequency Cepstral Coefficients\n(MFCCs).\n

Related