vix.ing · top · new · best · stats · spec

Monaural Speech Enhancement using Deep Neural Networks by Maximizing a\n Short-Time Objective Intelligibility Measure

2018/02/02 by Morten Kolbæk, Kolbæk, Morten, Zheng‐Hua Tan +3
Computer Science · Engineering · Neuroscience · #Advanced Adaptive Filtering Techniques #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Hearing Loss and Rehabilitation #Sound (cs.SD) #Speech and Audio Processing #Structural Health Monitoring Techniques #electronic engineering #information engineering

paper · pdf · doi:10.48550/arxiv.1802.00604

openalex publication_date 2018/02/02 · openalex created_date 2022/10/03 · openalex updated_date 2026/07/28

Abstract

In this paper we propose a Deep Neural Network (DNN) based Speech Enhancement\n(SE) system that is designed to maximize an approximation of the Short-Time\nObjective Intelligibility (STOI) measure. We formalize an approximate-STOI cost\nfunction and derive analytical expressions for the gradients required for DNN\ntraining and show that these gradients have desirable properties when used\ntogether with gradient based optimization techniques. We show through\nsimulation experiments that the proposed SE system achieves large improvements\nin estimated speech intelligibility, when tested on matched and unmatched\nnatural noise types, at multiple signal-to-noise ratios. Furthermore, we show\nthat the SE system, when trained using an approximate-STOI cost function\nperforms on par with a system trained with a mean square error cost applied to\nshort-time temporal envelopes. Finally, we show that the proposed SE system\nperforms on par with a traditional DNN based Short-Time Spectral Amplitude\n(STSA) SE system in terms of estimated speech intelligibility. These results\nare important because they suggest that traditional DNN based STSA SE systems\nmight be optimal in terms of estimated speech intelligibility.\n

Citations

Related