2021/12/01 by Hao Tan, Tan, Hao Hao
Computer Science · Neuroscience · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Music Technology and Sound Studies #Music and Audio Processing #Neuroscience and Music Perception #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
paper · pdf · doi:10.48550/arxiv.2112.00702
openalex publication_date 2021/12/01 · openalex created_date 2023/02/18 · openalex updated_date 2026/07/28
We present Mirable's submission to the 2021 Emotions and Themes in Music\nchallenge. In this work, we intend to address the question: can we leverage\nsemi-supervised learning techniques on music emotion recognition? With that, we\nexperiment with noisy student training, which has improved model performance in\nthe image classification domain. As the noisy student method requires a strong\nteacher model, we further delve into the factors including (i) input training\nlength and (ii) complementary music representations to further boost the\nperformance of the teacher model. For (i), we find that models trained with\nshort input length perform better in PR-AUC, whereas those trained with long\ninput length perform better in ROC-AUC. For (ii), we find that using harmonic\npitch class profiles (HPCP) consistently improve tagging performance, which\nsuggests that harmonic representation is useful for music emotion tagging.\nFinally, we find that noisy student method only improves tagging results for\nthe case of long training length. Additionally, we find that ensembling\nrepresentations trained with different training lengths can improve tagging\nresults significantly, which suggest a possible direction to explore\nincorporating multiple temporal resolutions in the network architecture for\nfuture work.\n