vix.ing · top · new · best · stats · spec

Does Phase Matter For Monaural Source Separation?

2017/11/02 by Mohit L. Dubey, Mohit Dubey, Dubey, Mohit +8
Computer Science · Engineering · Neuroscience · #Audio and Speech Processing (eess.AS) #Blind Source Separation Techniques #FOS: Computer and information sciences #FOS: Electrical engineering #Hearing Loss and Rehabilitation #Neural and Evolutionary Computing (cs.NE) #Sound (cs.SD) #Speech and Audio Processing #cs.NE #cs.SD #eess.AS #electronic engineering #information engineering

paper · pdf · doi:10.48550/arxiv.1711.00913

4 pages, 2 figures, NIPS format

arxiv created 2017/11/02 · openalex publication_date 2017/11/02 · arxiv updated 2017/11/06 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28

Abstract

The "cocktail party" problem of fully separating multiple sources from a single channel audio waveform remains unsolved. Current biological understanding of neural encoding suggests that phase information is preserved and utilized at every stage of the auditory pathway. However, current computational approaches primarily discard phase information in order to mask amplitude spectrograms of sound. In this paper, we seek to address whether preserving phase information in spectral representations of sound provides better results in monaural separation of vocals from a musical track by using a neurally plausible sparse generative model. Our results demonstrate that preserving phase information reduces artifacts in the separated tracks, as quantified by the signal to artifact ratio (GSAR). Furthermore, our proposed method achieves state-of-the-art performance for source separation, as quantified by a mean signal to interference ratio (GSIR) of 19.46.

Related