English

Toward Robust Real-World Audio Deepfake Detection: Closing the Explainability Gap

Machine Learning 2024-10-11 v1 Sound Audio and Speech Processing

Abstract

The rapid proliferation of AI-manipulated or generated audio deepfakes poses serious challenges to media integrity and election security. Current AI-driven detection solutions lack explainability and underperform in real-world settings. In this paper, we introduce novel explainability methods for state-of-the-art transformer-based audio deepfake detectors and open-source a novel benchmark for real-world generalizability. By narrowing the explainability gap between transformer-based audio deepfake detectors and traditional methods, our results not only build trust with human experts, but also pave the way for unlocking the potential of citizen intelligence to overcome the scalability issue in audio deepfake detection.

Keywords

Cite

@article{arxiv.2410.07436,
  title  = {Toward Robust Real-World Audio Deepfake Detection: Closing the Explainability Gap},
  author = {Georgia Channing and Juil Sock and Ronald Clark and Philip Torr and Christian Schroeder de Witt},
  journal= {arXiv preprint arXiv:2410.07436},
  year   = {2024}
}
R2 v1 2026-06-28T19:15:20.670Z