2020/10/22 by Woosung Choi, Minseok Kim, Choi, Woosung +5 · 2 citations
Computer Science · #Audio and Speech Processing (eess.AS) #Blind Source Separation Techniques #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
paper · pdf · doi:10.48550/arxiv.2010.11631
openalex publication_date 2020/10/22 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Recent deep-learning approaches have shown that Frequency Transformation (FT)\nblocks can significantly improve spectrogram-based single-source separation\nmodels by capturing frequency patterns. The goal of this paper is to extend the\nFT block to fit the multi-source task. We propose the Latent Source Attentive\nFrequency Transformation (LaSAFT) block to capture source-dependent frequency\npatterns. We also propose the Gated Point-wise Convolutional Modulation\n(GPoCM), an extension of Feature-wise Linear Modulation (FiLM), to modulate\ninternal features. By employing these two novel methods, we extend the\nConditioned-U-Net (CUNet) for multi-source separation, and the experimental\nresults indicate that our LaSAFT and GPoCM can improve the CUNet's performance,\nachieving state-of-the-art SDR performance on several MUSDB18 source separation\ntasks.\n