English

On the Role of Spatial Features in Foundation-Model-Based Speaker Diarization

Audio and Speech Processing 2026-01-06 v1

Abstract

Recent advances in speaker diarization exploit large pretrained foundation models, such as WavLM, to achieve state-of-the-art performance on multiple datasets. Systems like DiariZen leverage these rich single-channel representations, but are limited to single-channel audio, preventing the use of spatial cues available in multi-channel recordings. This work analyzes the impact of incorporating spatial information into a state-of-the-art single-channel diarization system by evaluating several strategies for conditioning the model on multi-channel spatial features. Experiments on meeting-style datasets indicate that spatial information can improve diarization performance, but the overall improvement is smaller than expected for the proposed system, suggesting that the features aggregated over all WavLM layers already capture much of the information needed for accurate speaker discrimination, also in overlapping speech regions. These findings provide insight into the potential and limitations of using spatial cues to enhance foundation model-based diarization.

Keywords

Cite

@article{arxiv.2601.02231,
  title  = {On the Role of Spatial Features in Foundation-Model-Based Speaker Diarization},
  author = {Marc Deegen and Tobias Gburrek and Tobias Cord-Landwehr and Thilo von Neumann and Jiangyu Han and Lukáš Burget and Reinhold Haeb-Umbach},
  journal= {arXiv preprint arXiv:2601.02231},
  year   = {2026}
}

Comments

Accepted at HSCMA 2026

R2 v1 2026-07-01T08:51:05.258Z