English

Pushing the Limits of End-to-End Diarization

Sound 2025-09-19 v1

Abstract

In this paper, we present state-of-the-art diarization error rates (DERs) on multiple publicly available datasets, including AliMeeting-far, AliMeeting-near, AMI-Mix, AMI-SDM, DIHARD III, and MagicData RAMC. Leveraging EEND-TA, a single unified non-autoregressive model for end-to-end speaker diarization, we achieve new benchmark results, most notably a DER of 14.49% on DIHARD III. Our approach scales pretraining through 8-speaker simulation mixtures, ensuring each generated speaker mixture configuration is sufficiently represented. These experiments highlight that EEND-based architectures possess a greater capacity for learning than previously explored, surpassing many existing diarization solutions while maintaining efficient speeds during inference.

Keywords

Cite

@article{arxiv.2509.14737,
  title  = {Pushing the Limits of End-to-End Diarization},
  author = {Samuel J. Broughton and Lahiru Samarakoon},
  journal= {arXiv preprint arXiv:2509.14737},
  year   = {2025}
}

Comments

As presented at Interspeech 2025

R2 v1 2026-07-01T05:43:23.596Z