2014/02/28 by Peter Foster, Matthias Mauch, Simon Dixon
Computer Science · Mathematics · #Artificial intelligence #Computer science #Correlation #Data mining #Image (mathematics) #Mathematics #Music Technology and Sound Studies #Music and Audio Processing #Pairwise comparison #Pattern recognition (psychology) #Rank (graph theory) #Similarity (geometry) #Speech and Audio Processing #cs.IR #cs.LG #cs.SD
paper · pdf · doi:10.1109/taslp.2014.2357676
published as IEEE/ACM Transactions on Audio, Speech, and Language Processing, vol. 22 no. 12, pp. 1965-1977, 2014 · 13 pages, 9 figures, 8 tables. Accepted version
openalex publication_date 2014/09/12 · arxiv created 2014/09/28 · arxiv updated 2014/09/30 · openalex created_date 2025/10/10 · openalex updated_date 2026/08/05
We propose string compressibility as a descriptor of temporal structure in audio, for the purpose of determining musical similarity. Our descriptors are based on computing track-wise compression rates of quantized audio features, using multiple temporal resolutions and quantization granularities. To verify that our descriptors capture musically relevant information, we incorporate our descriptors into similarity rating prediction and song year prediction tasks. We base our evaluation on a dataset of 15 500 track excerpts of Western popular music, for which we obtain 7 800 web-sourced pairwise similarity ratings. To assess the agreement among similarity ratings, we perform an evaluation under controlled conditions, obtaining a rank correlation of 0.33 between intersected sets of ratings. Combined with bag-of-features descriptors, we obtain performance gains of 31.1% and 10.9% for similarity rating prediction and song year prediction. For both tasks, analysis of selected descriptors reveals that representing features at multiple time scales benefits prediction accuracy.