vix.ing · top · new · best · stats · spec

Multi-Band Multi-Resolution Fully Convolutional Neural Networks for\n Singing Voice Separation

2019/10/21 by Emad M. Grais, Fei Zhao, Grais, Emad M. +3
Computer Science · #62H25 #68T01 #68T10 #68T45 #Audio and Speech Processing (eess.AS) #Blind Source Separation Techniques #FOS: Computer and information sciences #FOS: Electrical engineering #H.5.5 #I.2 #I.2.6 #I.4 #I.4.3 #I.5 #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Music and Audio Processing #Signal Processing (eess.SP) #Sound (cs.SD) #Speech and Audio Processing #electronic engineering #information engineering

paper · pdf · doi:10.48550/arxiv.1910.09266

openalex publication_date 2019/10/21 · openalex created_date 2022/07/28 · openalex updated_date 2026/07/28

Abstract

Deep neural networks with convolutional layers usually process the entire\nspectrogram of an audio signal with the same time-frequency resolutions, number\nof filters, and dimensionality reduction scale. According to the constant-Q\ntransform, good features can be extracted from audio signals if the low\nfrequency bands are processed with high frequency resolution filters and the\nhigh frequency bands with high time resolution filters. In the spectrogram of a\nmixture of singing voices and music signals, there is usually more information\nabout the voice in the low frequency bands than the high frequency bands. These\nraise the need for processing each part of the spectrogram differently. In this\npaper, we propose a multi-band multi-resolution fully convolutional neural\nnetwork (MBR-FCN) for singing voice separation. The MBR-FCN processes the\nfrequency bands that have more information about the target signals with more\nfilters and smaller dimentionality reduction scale than the bands with less\ninformation. Furthermore, the MBR-FCN processes the low frequency bands with\nhigh frequency resolution filters and the high frequency bands with high time\nresolution filters. Our experimental results show that the proposed MBR-FCN\nwith very few parameters achieves better singing voice separation performance\nthan other deep neural networks.\n

Related