English

Unmute the Patch Tokens: Rethinking Probing in Multi-Label Audio Classification

Sound 2025-12-04 v3 Machine Learning

Abstract

Although probing frozen models has become a standard evaluation paradigm, self-supervised learning in audio defaults to fine-tuning when pursuing state-of-the-art on AudioSet. A key reason is that global pooling creates an information bottleneck causing linear probes to misrepresent the embedding quality: The cls\texttt{cls}-token discards crucial token information about dispersed, localized events in audio. This weakness is rooted in the mismatch between the pretraining objective (globally) and the downstream task (localized). Across a comprehensive benchmark of 13 datasets and 6 spectrogram-based encoders, we investigate the global pooling bottleneck. We introduce binarized prototypical probes: a lightweight and simple pooling method that learns prototypes to perform class-wise information aggregation. Despite its simplicity, our method notably outperforms linear and attentive probing. Our work establishes probing as a competitive and efficient paradigm for evaluating audio SSL models, challenging the reliance on costly fine-tuning.

Keywords

Cite

@article{arxiv.2509.24901,
  title  = {Unmute the Patch Tokens: Rethinking Probing in Multi-Label Audio Classification},
  author = {Lukas Rauch and René Heinrich and Houtan Ghaffari and Lukas Miklautz and Ilyass Moummad and Bernhard Sick and Christoph Scholz},
  journal= {arXiv preprint arXiv:2509.24901},
  year   = {2025}
}

Comments

Currently under review

R2 v1 2026-07-01T06:04:47.764Z