vix.ing · top · new · best · stats · spec

Continuous Modeling of the Denoising Process for Speech Enhancement Based on Deep Learning

2023/09/17 by Zilu Guo, Jun Du, Guo, Zilu +3
Computer Science · Engineering · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #Ultrasonics and Acoustic Wave Propagation #electronic engineering #information engineering

paper · pdf · doi:10.48550/arxiv.2309.09270

openalex publication_date 2023/09/17 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28

Abstract

In this paper, we explore a continuous modeling approach for deep-learning-based speech enhancement, focusing on the denoising process. We use a state variable to indicate the denoising process. The starting state is noisy speech and the ending state is clean speech. The noise component in the state variable decreases with the change of the state index until the noise component is 0. During training, a UNet-like neural network learns to estimate every state variable sampled from the continuous denoising process. In testing, we introduce a controlling factor as an embedding, ranging from zero to one, to the neural network, allowing us to control the level of noise reduction. This approach enables controllable speech enhancement and is adaptable to various application scenarios. Experimental results indicate that preserving a small amount of noise in the clean target benefits speech enhancement, as evidenced by improvements in both objective speech measures and automatic speech recognition performance.

Related