2019/11/11 by Minz Won, Won, Minz, Sanghyuk Chun +3
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Image Retrieval and Classification Techniques #Music and Audio Processing #Sound (cs.SD) #Video Analysis and Summarization #electronic engineering #information engineering
paper · pdf · doi:10.48550/arxiv.1911.04385
openalex publication_date 2019/11/11 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Recently, we proposed a self-attention based music tagging model. Different from most of the conventional deep architectures in music information retrieval, which use stacked 3x3 filters by treating music spectrograms as images, the proposed self-attention based model attempted to regard music as a temporal sequence of individual audio events. Not only the performance, but it could also facilitate better interpretability. In this paper, we mainly focus on visualizing and understanding the proposed self-attention based music tagging model.