2019/10/22 by Ethan Manilow, Prem Seetharaman, Manilow, Ethan +3 · 1 citation
Computer Science · #Speech and Audio Processing #Music and Audio Processing #Speech Recognition and Synthesis
paper · pdf · doi:10.48550/arxiv.1910.12621
We present a single deep learning architecture that can both separate an\naudio recording of a musical mixture into constituent single-instrument\nrecordings and transcribe these instruments into a human-readable format at the\nsame time, learning a shared musical representation for both tasks. This novel\narchitecture, which we call Cerberus, builds on the Chimera network for source\nseparation by adding a third "head" for transcription. By training each head\nwith different losses, we are able to jointly learn how to separate and\ntranscribe up to 5 instruments in our experiments with a single network. We\nshow that the two tasks are highly complementary with one another and when\nlearned jointly, lead to Cerberus networks that are better at both separation\nand transcription and generalize better to unseen mixtures.\n