vix.ing · top · new · best · stats · spec

Speech Enhancement with Wide Residual Networks in Reverberant\n Environments

2019/04/09 by Jorge Llombart, Llombart, Jorge, Dayana Ribas +9 · 1 citation
Computer Science · Engineering · Neuroscience · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Hearing Loss and Rehabilitation #Indoor and Outdoor Localization Technologies #Sound (cs.SD) #Speech and Audio Processing #electronic engineering #information engineering

paper · pdf · doi:10.48550/arxiv.1904.05167

openalex publication_date 2019/04/09 · openalex created_date 2022/07/24 · openalex updated_date 2026/07/28

Abstract

This paper proposes a speech enhancement method which exploits the high\npotential of residual connections in a Wide Residual Network architecture. This\nis supported on single dimensional convolutions computed alongside the time\ndomain, which is a powerful approach to process contextually correlated\nrepresentations through the temporal domain, such as speech feature sequences.\nWe find the residual mechanism extremely useful for the enhancement task since\nthe signal always has a linear shortcut and the non-linear path enhances it in\nseveral steps by adding or subtracting corrections. The enhancement capability\nof the proposal is assessed by objective quality metrics evaluated with\nsimulated and real samples of reverberated speech signals. Results show that\nthe proposal outperforms the state-of-the-art method called WPE, which is known\nto effectively reduce reverberation and greatly enhance the signal. The\nproposed model, trained with artificial synthesized reverberation data, was\nable to generalize to real room impulse responses for a variety of conditions\n(e.g. different room sizes, RT60, near & far field). Furthermore, it\nachieves accuracy for real speech with reverberation from two different\ndatasets.\n

Cited by

Related