2025/01/27 by Wang, Yueguan, Matsushima, Tatsunari, Soichiro Matsushima +2
Health Professions · Neuroscience · #Artificial Intelligence in Healthcare #Audio and Speech Processing (eess.AS) #Brain Tumor Detection and Classification #Computation and Language (cs.CL) #EEG and Brain-Computer Interfaces #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
paper · pdf · doi:10.48550/arxiv.2501.16201
openalex publication_date 2025/01/27 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
This study explores a multi-lingual audio self-supervised learning model for detecting mild cognitive impairment (MCI) using the TAUKADIAL cross-lingual dataset. While speech transcription-based detection with BERT models is effective, limitations exist due to a lack of transcriptions and temporal information. To address these issues, the study utilizes features directly from speech utterances with W2V-BERT-2.0. We propose a visualization method to detect essential layers of the model for MCI classification and design a specific inference logic considering the characteristics of MCI. The experiment shows competitive results, and the proposed inference logic significantly contributes to the improvements from the baseline. We also conduct detailed analysis which reveals the challenges related to speaker bias in the features and the sensitivity of MCI classification accuracy to the data split, providing valuable insights for future research.