2019/04/23 by Rohit Voleti, Stephanie M. Woolridge, Voleti, Rohit +9
Computer Science · Psychology · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Stuttering Research and Treatment #Text Readability and Simplification
paper · pdf · doi:10.48550/arxiv.1904.10622
openalex publication_date 2019/04/23 · openalex created_date 2022/07/24 · openalex updated_date 2026/07/28
Several studies have shown that speech and language features, automatically\nextracted from clinical interviews or spontaneous discourse, have diagnostic\nvalue for mental disorders such as schizophrenia and bipolar disorder. They\ntypically make use of a large feature set to train a classifier for\ndistinguishing between two groups of interest, i.e. a clinical and control\ngroup. However, a purely data-driven approach runs the risk of overfitting to a\nparticular data set, especially when sample sizes are limited. Here, we first\ndown-select the set of language features to a small subset that is related to a\nwell-validated test of functional ability, the Social Skills Performance\nAssessment (SSPA). This helps establish the concurrent validity of the selected\nfeatures. We use only these features to train a simple classifier to\ndistinguish between groups of interest. Linear regression reveals that a subset\nof language features can effectively model the SSPA, with a correlation\ncoefficient of 0.75. Furthermore, the same feature set can be used to build a\nstrong binary classifier to distinguish between healthy controls and a clinical\ngroup (AUC = 0.96) and also between patients within the clinical group with\nschizophrenia and bipolar I disorder (AUC = 0.83).\n