vix.ing · top · new · best · stats · spec

A consolidated view of loss functions for supervised deep learning-based\n speech enhancement

2020/09/25 by Sebastian Braun, Braun, Sebastian, Ivan Tashev +1 · 3 citations
Computer Science · Engineering · #Advanced Adaptive Filtering Techniques #Audio and Speech Processing (eess.AS) #FOS: Electrical engineering #Indoor and Outdoor Localization Technologies #Speech and Audio Processing #electronic engineering #information engineering

paper · pdf · doi:10.48550/arxiv.2009.12286

openalex publication_date 2020/09/25 · openalex created_date 2022/07/25 · openalex updated_date 2026/07/28

Abstract

Deep learning-based speech enhancement for real-time applications recently\nmade large advancements. Due to the lack of a tractable perceptual optimization\ntarget, many myths around training losses emerged, whereas the contribution to\nsuccess of the loss functions in many cases has not been investigated isolated\nfrom other factors such as network architecture, features, or training\nprocedures. In this work, we investigate a wide variety of loss spectral\nfunctions for a recurrent neural network architecture suitable to operate in\nonline frame-by-frame processing. We relate magnitude-only with phase-aware\nlosses, ratios, correlation metrics, and compressed metrics. Our results reveal\nthat combining magnitude-only with phase-aware objectives always leads to\nimprovements, even when the phase is not enhanced. Furthermore, using\ncompressed spectral values also yields a significant improvement. On the other\nhand, phase-sensitive improvement is best achieved by linear domain losses such\nas mean absolute error.\n

Cited by

Related