English

Categorical Unsupervised Variational Acoustic Clustering

Audio and Speech Processing 2026-01-22 v3

Abstract

We propose a categorical approach for unsupervised variational acoustic clustering of audio data in the time-frequency domain. The consideration of a categorical distribution enforces sharper clustering even when data points strongly overlap in time and frequency, which is the case for most datasets of urban acoustic scenes. To this end, we use a Gumbel-Softmax distribution as a soft approximation to the categorical distribution, allowing for training via backpropagation. In this settings, the softmax temperature serves as the main mechanism to tune clustering performance. The results show that the proposed model can obtain impressive clustering performance for all considered datasets, even when data points strongly overlap in time and frequency.

Keywords

Cite

@article{arxiv.2504.07652,
  title  = {Categorical Unsupervised Variational Acoustic Clustering},
  author = {Luan Vinícius Fiorio and Ivana Nikoloska and Ronald M. Aarts},
  journal= {arXiv preprint arXiv:2504.07652},
  year   = {2026}
}

Comments

Please refer to arXiv:2510.01940 for an extended version

R2 v1 2026-06-28T22:53:31.473Z