vix.ing · top · new · best · stats · spec

Similarity measures for vocal-based drum sample retrieval using deep\n convolutional auto-encoders

2018/02/14 by Adib Mehrabi, Mehrabi, Adib, Keunwoo Choi +5
Computer Science · #Music and Audio Processing #Speech and Audio Processing #Music Technology and Sound Studies

paper · pdf · doi:10.48550/arxiv.1802.05178

Abstract

The expressive nature of the voice provides a powerful medium for\ncommunicating sonic ideas, motivating recent research on methods for query by\nvocalisation. Meanwhile, deep learning methods have demonstrated\nstate-of-the-art results for matching vocal imitations to imitated sounds, yet\nlittle is known about how well learned features represent the perceptual\nsimilarity between vocalisations and queried sounds. In this paper, we address\nthis question using similarity ratings between vocal imitations and imitated\ndrum sounds. We use a linear mixed effect regression model to show how features\nlearned by convolutional auto-encoders (CAEs) perform as predictors for\nperceptual similarity between sounds. Our experiments show that CAEs outperform\nthree baseline feature sets (spectrogram-based representations, MFCCs, and\ntemporal features) at predicting the subjective similarity ratings. We also\ninvestigate how the size and shape of the encoded layer effects the predictive\npower of the learned features. The results show that preservation of temporal\ninformation is more important than spectral resolution for this application.\n

Related