vix.ing · top · new · best · stats · spec

Unsupervised Interpretable Representation Learning for Singing Voice\n Separation

2020/03/03 by Stylianos Ioannis Mimilakis, Mimilakis, Stylianos I., Konstantinos Drossos +3
Computer Science · #Speech and Audio Processing #Music and Audio Processing #Speech Recognition and Synthesis

paper · pdf · doi:10.48550/arxiv.2003.01567

Abstract

In this work, we present a method for learning interpretable music signal\nrepresentations directly from waveform signals. Our method can be trained using\nunsupervised objectives and relies on the denoising auto-encoder model that\nuses a simple sinusoidal model as decoding functions to reconstruct the singing\nvoice. To demonstrate the benefits of our method, we employ the obtained\nrepresentations to the task of informed singing voice separation via binary\nmasking, and measure the obtained separation quality by means of\nscale-invariant signal to distortion ratio. Our findings suggest that our\nmethod is capable of learning meaningful representations for singing voice\nseparation, while preserving conveniences of the the short-time Fourier\ntransform like non-negativity, smoothness, and reconstruction subject to\ntime-frequency masking, that are desired in audio and music source separation.\n

Related