vix.ing · top · new · best · stats · spec

Improved Intelligibility of Dysarthric Speech using Conditional Flow Matching

2025/06/19 by Das, Shoutrik, Singh, Nishant, Gangwar, Arjun +1
Computer Science · Medicine · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Phonocardiography and Auscultation Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Voice and Speech Disorders #electronic engineering #information engineering

paper · pdf · doi:10.48550/arxiv.2506.16127

openalex publication_date 2025/06/19 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28

Abstract

Dysarthria is a neurological disorder that significantly impairs speech intelligibility, often rendering affected individuals unable to communicate effectively. This necessitates the development of robust dysarthric-to-regular speech conversion techniques. In this work, we investigate the utility and limitations of self-supervised learning (SSL) features and their quantized representations as an alternative to mel-spectrograms for speech generation. Additionally, we explore methods to mitigate speaker variability by generating clean speech in a single-speaker voice using features extracted from WavLM. To this end, we propose a fully non-autoregressive approach that leverages Conditional Flow Matching (CFM) with Diffusion Transformers to learn a direct mapping from dysarthric to clean speech. Our findings highlight the effectiveness of discrete acoustic units in improving intelligibility while achieving faster convergence compared to traditional mel-spectrogram-based approaches.

Citations

Related