2019/10/28 by Khaled Koutini, Koutini, Khaled, Shreyan Chowdhury +7
Computer Science · #Music and Audio Processing #Music Technology and Sound Studies #Speech and Audio Processing
paper · pdf · doi:10.48550/arxiv.1911.05833
We present CP-JKU submission to MediaEval 2019; a Receptive Field-(RF)-regularized and Frequency-Aware CNN approach for tagging music with emotion/mood labels. We perform an investigation regarding the impact of the RF of the CNNs on their performance on this dataset. We observe that ResNets with smaller receptive fields -- originally adapted for acoustic scene classification -- also perform well in the emotion tagging task. We improve the performance of such architectures using techniques such as Frequency Awareness and Shake-Shake regularization, which were used in previous work on general acoustic recognition tasks.