2025/09/29 by Lukas Rauch, René Heinrich, Rauch, Lukas +11 · 2 citations
Arts and Humanities · Computer Science · Engineering · #Diverse Musicological Studies #FOS: Computer and information sciences #Machine Learning (cs.LG) #Music and Audio Processing #Sound (cs.SD) #Water Systems and Optimization
paper · pdf · doi:10.48550/arxiv.2509.24901
openalex publication_date 2025/09/29 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Although probing frozen models has become a standard evaluation paradigm, self-supervised learning in audio defaults to fine-tuning when pursuing state-of-the-art on AudioSet. A key reason is that global pooling creates an information bottleneck causing linear probes to misrepresent the embedding quality: The cls-token discards crucial token information about dispersed, localized events in audio. This weakness is rooted in the mismatch between the pretraining objective (globally) and the downstream task (localized). Across a comprehensive benchmark of 13 datasets and 6 spectrogram-based encoders, we investigate the global pooling bottleneck. We introduce binarized prototypical probes: a lightweight and simple pooling method that learns prototypes to perform class-wise information aggregation. Despite its simplicity, our method notably outperforms linear and attentive probing. Our work establishes probing as a competitive and efficient paradigm for evaluating audio SSL models, challenging the reliance on costly fine-tuning.