English

Face-Voice Association for Audiovisual Active Speaker Detection in Egocentric Recordings

Multimedia 2025-06-24 v1 Sound Audio and Speech Processing

Abstract

Audiovisual active speaker detection (ASD) is conventionally performed by modelling the temporal synchronisation of acoustic and visual speech cues. In egocentric recordings, however, the efficacy of synchronisation-based methods is compromised by occlusions, motion blur, and adverse acoustic conditions. In this work, a novel framework is proposed that exclusively leverages cross-modal face-voice associations to determine speaker activity. An existing face-voice association model is integrated with a transformer-based encoder that aggregates facial identity information by dynamically weighting each frame based on its visual quality. This system is then coupled with a front-end utterance segmentation method, producing a complete ASD system. This work demonstrates that the proposed system, Self-Lifting for audiovisual active speaker detection(SL-ASD), achieves performance comparable to, and in certain cases exceeding, that of parameter-intensive synchronisation-based approaches with significantly fewer learnable parameters, thereby validating the feasibility of substituting strict audiovisual synchronisation modelling with flexible biometric associations in challenging egocentric scenarios.

Keywords

Cite

@article{arxiv.2506.18055,
  title  = {Face-Voice Association for Audiovisual Active Speaker Detection in Egocentric Recordings},
  author = {Jason Clarke and Yoshihiko Gotoh and Stefan Goetze},
  journal= {arXiv preprint arXiv:2506.18055},
  year   = {2025}
}

Comments

Accepted to EUSIPCO 2025. 5 pages, 1 figure. To appear in the Proceedings of the 33rd European Signal Processing Conference (EUSIPCO), September 8-12, 2025, Palermo, Italy