vix.ing · top · new · best · stats

DCCRGAN: Deep Complex Convolution Recurrent Generator Adversarial Network for Speech Enhancement

2020/12/19 by Huixiang Huang, Renjie Wu, Huang, Huixiang +7 · 2 citations
Computer Science · Engineering · #Advanced Adaptive Filtering Techniques #Algorithm #Artificial intelligence #Artificial neural network #Computer engineering #Computer science #Convolution (computer science) #Deep learning #Engineering #Generative adversarial network #Generator (circuit theory) #Image and Signal Denoising Methods #Noise reduction #Power (physics) #Speech and Audio Processing #Speech enhancement #Speech recognition #Task (project management) #Telecommunications #Waveform #cs.SD #eess.AS

paper · pdf · doi:10.48550/arxiv.2012.10732

published in arXiv (Cornell University) (Cornell University)

openalex publication_date 2020/12/19 · arxiv created 2021/03/07 · arxiv updated 2021/03/09 · openalex created_date 2022/07/25 · openalex updated_date 2026/07/28

Abstract

Generative adversarial network (GAN) still exists some problems in dealing with speech enhancement (SE) task. Some GAN-based systems adopt the same structure from Pixel-to-Pixel directly without special optimization. The importance of the generator network has not been fully explored. Other related researches change the generator network but operate in the time-frequency domain, which ignores the phase mismatch problem. In order to solve these problems, a deep complex convolution recurrent GAN (DCCRGAN) structure is proposed in this paper. The complex module builds the correlation between magnitude and phase of the waveform and has been proved to be effective. The proposed structure is trained in an end-to-end way. Different LSTM layers are used in the generator network to sufficiently explore the speech enhancement performance of DCCRGAN. The experimental results confirm that the proposed DCCRGAN outperforms the state-of-the-art GAN-based SE systems.

Related