Frame-wise and overlap-robust speaker embeddings for meeting diarization
Audio and Speech Processing
2023-06-02 v1
Abstract
Using a Teacher-Student training approach we developed a speaker embedding extraction system that outputs embeddings at frame rate. Given this high temporal resolution and the fact that the student produces sensible speaker embeddings even for segments with speech overlap, the frame-wise embeddings serve as an appropriate representation of the input speech signal for an end-to-end neural meeting diarization (EEND) system. We show in experiments that this representation helps mitigate a well-known problem of EEND systems: when increasing the number of speakers the diarization performance drop is significantly reduced. We also introduce block-wise processing to be able to diarize arbitrarily long meetings.
Cite
@article{arxiv.2306.00625,
title = {Frame-wise and overlap-robust speaker embeddings for meeting diarization},
author = {Tobias Cord-Landwehr and Christoph Boeddeker and Cătălin Zorilă and Rama Doddipatla and Reinhold Haeb-Umbach},
journal= {arXiv preprint arXiv:2306.00625},
year = {2023}
}
Comments
ICASSP 2023