English

Quantifying the Uncertainty of Blindly Estimated Room Embeddings Using a Dispersion-Calibrated Score

Sound 2026-07-01 v1 Machine Learning

Abstract

Room embeddings derived from reverberant speech are often unreliable: speech content and recording degradation can alter the representation even when speaker, room, and source-receiver geometry remain unchanged, degrading downstream task performance. We propose a framework that learns room embeddings robust to speech-content variation and a representation-level uncertainty score from reverberant speech without downstream-task supervision. The embedding is anchored to a structured room impulse response (RIR) latent space and trained using a multi-view data structure with Kullback-Leibler (KL)-based alignment; a multi-positive contrastive term further refines robustness. A lightweight uncertainty head is calibrated using the dispersion of corruption-induced embeddings and optimized with a rank-based objective. Across waveform- and spectrogram-level corruptions, the score is consistent with representation dispersion and enables effective selective prediction while requiring only a single utterance at inference.

Keywords

Cite

@article{arxiv.2607.01527,
  title  = {Quantifying the Uncertainty of Blindly Estimated Room Embeddings Using a Dispersion-Calibrated Score},
  author = {Yang Xiang and Philipp Götz and Emanuël A. P. Habets and Andreas Walther and Wenwu Wang and Philip J. B. Jackson},
  journal= {arXiv preprint arXiv:2607.01527},
  year   = {2026}
}

Comments

Accepted to INTERSPEECH 2026