2020/10/22 by Yun-Ning Hung, Gordon Wichern, Hung, Yun-Ning +3
Computer Science · #Speech and Audio Processing #Music and Audio Processing #Speech Recognition and Synthesis
paper · pdf · doi:10.48550/arxiv.2010.11904
Most music source separation systems require large collections of isolated\nsources for training, which can be difficult to obtain. In this work, we use\nmusical scores, which are comparatively easy to obtain, as a weak label for\ntraining a source separation system. In contrast with previous score-informed\nseparation approaches, our system does not require isolated sources, and score\nis used only as a training target, not required for inference. Our model\nconsists of a separator that outputs a time-frequency mask for each instrument,\nand a transcriptor that acts as a critic, providing both temporal and frequency\nsupervision to guide the learning of the separator. A harmonic mask constraint\nis introduced as another way of leveraging score information during training,\nand we propose two novel adversarial losses for additional fine-tuning of both\nthe transcriptor and the separator. Results demonstrate that using score\ninformation outperforms temporal weak-labels, and adversarial structures lead\nto further improvements in both separation and transcription performance.\n