English

Rehearsal with Auxiliary-Informed Sampling for Audio Deepfake Detection

Sound 2026-01-22 v1 Artificial Intelligence Cryptography and Security Machine Learning Audio and Speech Processing

Abstract

The performance of existing audio deepfake detection frameworks degrades when confronted with new deepfake attacks. Rehearsal-based continual learning (CL), which updates models using a limited set of old data samples, helps preserve prior knowledge while incorporating new information. However, existing rehearsal techniques don't effectively capture the diversity of audio characteristics, introducing bias and increasing the risk of forgetting. To address this challenge, we propose Rehearsal with Auxiliary-Informed Sampling (RAIS), a rehearsal-based CL approach for audio deepfake detection. RAIS employs a label generation network to produce auxiliary labels, guiding diverse sample selection for the memory buffer. Extensive experiments show RAIS outperforms state-of-the-art methods, achieving an average Equal Error Rate (EER) of 1.953 % across five experiences. The code is available at: https://github.com/falihgoz/RAIS.

Keywords

Cite

@article{arxiv.2505.24486,
  title  = {Rehearsal with Auxiliary-Informed Sampling for Audio Deepfake Detection},
  author = {Falih Gozi Febrinanto and Kristen Moore and Chandra Thapa and Jiangang Ma and Vidya Saikrishna and Feng Xia},
  journal= {arXiv preprint arXiv:2505.24486},
  year   = {2026}
}

Comments

Accepted by Interspeech 2025

R2 v1 2026-07-01T02:50:26.189Z