vix.ing · top · new · best · stats · spec

Efficient Trainable Front-Ends for Neural Speech Enhancement

2020/02/20 by Jonah Casebeer, Casebeer, Jonah, Umut Isik +5
Computer Science · Engineering · Mathematics · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Neural and Evolutionary Computing (cs.NE) #Sound (cs.SD) #cs.LG #cs.NE #cs.SD #eess.AS #electronic engineering #information engineering #stat.ML

paper · pdf · doi:10.48550/arxiv.2002.09286

5 pages, 5 figures, ICASSP 2020

arxiv created 2020/02/20 · arxiv updated 2020/02/24

Abstract

Many neural speech enhancement and source separation systems operate in the time-frequency domain. Such models often benefit from making their Short-Time Fourier Transform (STFT) front-ends trainable. In current literature, these are implemented as large Discrete Fourier Transform matrices; which are prohibitively inefficient for low-compute systems. We present an efficient, trainable front-end based on the butterfly mechanism to compute the Fast Fourier Transform, and show its accuracy and efficiency benefits for low-compute neural speech enhancement models. We also explore the effects of making the STFT window trainable.

Related