English

DiaPer: End-to-End Neural Diarization with Perceiver-Based Attractors

Audio and Speech Processing 2024-06-04 v3 Sound

Abstract

Until recently, the field of speaker diarization was dominated by cascaded systems. Due to their limitations, mainly regarding overlapped speech and cumbersome pipelines, end-to-end models have gained great popularity lately. One of the most successful models is end-to-end neural diarization with encoder-decoder based attractors (EEND-EDA). In this work, we replace the EDA module with a Perceiver-based one and show its advantages over EEND-EDA; namely obtaining better performance on the largely studied Callhome dataset, finding the quantity of speakers in a conversation more accurately, and faster inference time. Furthermore, when exhaustively compared with other methods, our model, DiaPer, reaches remarkable performance with a very lightweight design. Besides, we perform comparisons with other works and a cascaded baseline across more than ten public wide-band datasets. Together with this publication, we release the code of DiaPer as well as models trained on public and free data.

Keywords

Cite

@article{arxiv.2312.04324,
  title  = {DiaPer: End-to-End Neural Diarization with Perceiver-Based Attractors},
  author = {Federico Landini and Mireia Diez and Themos Stafylakis and Lukáš Burget},
  journal= {arXiv preprint arXiv:2312.04324},
  year   = {2024}
}

Comments

Accepted by IEEE/ACM Transactions on Audio, Speech, and Language Processing

R2 v1 2026-06-28T13:44:01.210Z