vix.ing · top · new · best · stats · spec

Exploring Generative Error Correction for Dysarthric Speech Recognition

2025/05/26 by Moreno La Quatra, La Quatra, Moreno, Alkis Koudounas +5 · 2 citations
Computer Science · Medicine · Psychology · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Phonetics and Phonology Research #Speech Recognition and Synthesis #Voice and Speech Disorders #electronic engineering #information engineering

paper · pdf · doi:10.48550/arxiv.2505.20163

openalex publication_date 2025/05/26 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28

Abstract

Despite the remarkable progress in end-to-end Automatic Speech Recognition (ASR) engines, accurately transcribing dysarthric speech remains a major challenge. In this work, we proposed a two-stage framework for the Speech Accessibility Project Challenge at INTERSPEECH 2025, which combines cutting-edge speech recognition models with LLM-based generative error correction (GER). We assess different configurations of model scales and training strategies, incorporating specific hypothesis selection to improve transcription accuracy. Experiments on the Speech Accessibility Project dataset demonstrate the strength of our approach on structured and spontaneous speech, while highlighting challenges in single-word recognition. Through comprehensive analysis, we provide insights into the complementary roles of acoustic and linguistic modeling in dysarthric speech recognition

Citations

Cited by

Related