vix.ing · top · new · best · stats · spec

Whispered-to-voiced Alaryngeal Speech Conversion with Generative\n Adversarial Networks

2018/08/31 by Santiago Pascual, Pascual, Santiago, Antonio Bonafonte +5
Computer Science · Medicine · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #Voice and Speech Disorders #electronic engineering #information engineering

paper · pdf · doi:10.48550/arxiv.1808.10687

openalex publication_date 2018/08/31 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28

Abstract

Most methods of voice restoration for patients suffering from aphonia either\nproduce whispered or monotone speech. Apart from intelligibility, this type of\nspeech lacks expressiveness and naturalness due to the absence of pitch\n(whispered speech) or artificial generation of it (monotone speech). Existing\ntechniques to restore prosodic information typically combine a vocoder, which\nparameterises the speech signal, with machine learning techniques that predict\nprosodic information. In contrast, this paper describes an end-to-end neural\napproach for estimating a fully-voiced speech waveform from whispered\nalaryngeal speech. By adapting our previous work in speech enhancement with\ngenerative adversarial networks, we develop a speaker-dependent model to\nperform whispered-to-voiced speech conversion. Preliminary qualitative results\nshow effectiveness in re-generating voiced speech, with the creation of\nrealistic pitch contours.\n

Related