2024/07/03 by Dominika Woszczyk, Woszczyk, Dominika, Ranya Aloufi +3 · 1 citation
Computer Science · Psychology · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #Cryptography and Security (cs.CR) #Digital Mental Health Interventions #Emotion and Mood Recognition #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Sound (cs.SD) #User Authentication and Security Systems #electronic engineering #information engineering
paper · pdf · doi:10.48550/arxiv.2407.03470
openalex publication_date 2024/07/03 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Speaker embeddings extracted from voice recordings have been proven valuable for dementia detection. However, by their nature, these embeddings contain identifiable information which raises privacy concerns. In this work, we aim to anonymize embeddings while preserving the diagnostic utility for dementia detection. Previous studies rely on adversarial learning and models trained on the target attribute and struggle in limited-resource settings. We propose a novel approach that leverages domain knowledge to disentangle prosody features relevant to dementia from speaker embeddings without relying on a dementia classifier. Our experiments show the effectiveness of our approach in preserving speaker privacy (speaker recognition F1-score .01%) while maintaining high dementia detection score F1-score of 74% on the ADReSS dataset. Our results are also on par with a more constrained classifier-dependent system on ADReSSo (.01% and .66%), and have no impact on synthesized speech naturalness.