Open-Set Source Tracing of Audio Deepfake Systems
Abstract
Existing research on source tracing of audio deepfake systems has focused primarily on the closed-set scenario, while studies that evaluate open-set performance are limited to a small number of unseen systems. Due to the large number of emerging audio deepfake systems, robust open-set source tracing is critical. We leverage the protocol of the Interspeech 2025 special session on source tracing to evaluate methods for improving open-set source tracing performance. We introduce a novel adaptation to the energy score for out-of-distribution (OOD) detection, softmax energy (SME). We find that replacing the typical temperature-scaled energy score with SME provides a relative average improvement of 31% in the standard FPR95 (false positive rate at true positive rate of 95%) measure. We further explore SME-guided training as well as copy synthesis, codec, and reverberation augmentations, yielding an FPR95 of 8.3%.
Cite
@article{arxiv.2507.06470,
title = {Open-Set Source Tracing of Audio Deepfake Systems},
author = {Nicholas Klein and Hemlata Tak and Elie Khoury},
journal= {arXiv preprint arXiv:2507.06470},
year = {2025}
}
Comments
Accepted by INTERSPEECH 2025 as part of the special session "Source Tracing: The Origins of Synthetic or Manipulated Speech"