English

Token-based Attractors and Cross-attention in Spoof Diarization

Audio and Speech Processing 2025-09-17 v1

Abstract

Spoof diarization identifies ``what spoofed when" in a given speech by temporally locating spoofed regions and determining their manipulation techniques. As a first step toward this task, prior work proposed a two-branch model for localization and spoof type clustering, which laid the foundation for spoof diarization. However, its simple structure limits the ability to capture complex spoofing patterns and lacks explicit reference points for distinguishing between bona fide and various spoofing types. To address these limitations, our approach introduces learnable tokens where each token represents acoustic features of bona fide and spoofed speech. These attractors interact with frame-level embeddings to extract discriminative representations, improving separation between genuine and generated speech. Vast experiments on PartialSpoof dataset consistently demonstrate that our approach outperforms existing methods in bona fide detection and spoofing method clustering.

Keywords

Cite

@article{arxiv.2509.13085,
  title  = {Token-based Attractors and Cross-attention in Spoof Diarization},
  author = {Kyo-Won Koo and Chan-yeong Lim and Jee-weon Jung and Hye-jin Shim and Ha-Jin Yu},
  journal= {arXiv preprint arXiv:2509.13085},
  year   = {2025}
}

Comments

Accepted to IEEE ASRU 2025

R2 v1 2026-07-01T05:39:27.084Z