English

Simultaneous Diarization and Separation of Meetings through the Integration of Statistical Mixture Models

Audio and Speech Processing 2025-02-25 v2

Abstract

We propose an approach for simultaneous diarization and separation of meeting data. It consists of a complex Angular Central Gaussian Mixture Model (cACGMM) for speech source separation, and a von-Mises-Fisher Mixture Model (VMFMM) for diarization in a joint statistical framework. Through the integration, both spatial and spectral information are exploited for diarization and separation. We also develop a method for counting the number of active speakers in a segment of a meeting to support block-wise processing. While the total number of speakers in a meeting may be known, it is usually not known on a per-segment level. With the proposed speaker counting, joint diarization and source separation can be done segment-by-segment, and the permutation problem across segments is solved, thus allowing for block-online processing in the future. Experimental results on the LibriCSS meeting corpus show that the integrated approach outperforms a cascaded approach of diarization and speech enhancement in terms of WER, both on a per-segment and on a per-meeting level.

Keywords

Cite

@article{arxiv.2410.21455,
  title  = {Simultaneous Diarization and Separation of Meetings through the Integration of Statistical Mixture Models},
  author = {Tobias Cord-Landwehr and Christoph Boeddeker and Reinhold Haeb-Umbach},
  journal= {arXiv preprint arXiv:2410.21455},
  year   = {2025}
}

Comments

Accepted at ICASSP2025