English

Room Impulse Responses help attackers to evade Deep Fake Detection

Audio and Speech Processing 2024-09-24 v1 Sound

Abstract

The ASVspoof 2021 benchmark, a widely-used evaluation framework for anti-spoofing, consists of two subsets: Logical Access (LA) and Deepfake (DF), featuring samples with varied coding characteristics and compression artifacts. Notably, the current state-of-the-art (SOTA) system boasts impressive performance, achieving an Equal Error Rate (EER) of 0.87% on the LA subset and 2.58% on the DF. However, benchmark accuracy is no guarantee of robustness in real-world scenarios. This paper investigates the effectiveness of utilizing room impulse responses (RIRs) to enhance fake speech and increase their likelihood of evading fake speech detection systems. Our findings reveal that this simple approach significantly improves the evasion rate, doubling the SOTA system's EER. To counter this type of attack, We augmented training data with a large-scale synthetic/simulated RIR dataset. The results demonstrate significant improvement on both reverberated fake speech and original samples, reducing DF task EER to 2.13%.

Keywords

Cite

@article{arxiv.2409.14712,
  title  = {Room Impulse Responses help attackers to evade Deep Fake Detection},
  author = {Hieu-Thi Luong and Duc-Tuan Truong and Kong Aik Lee and Eng Siong Chng},
  journal= {arXiv preprint arXiv:2409.14712},
  year   = {2024}
}

Comments

7 pages, to be presented at SLT 2024

R2 v1 2026-06-28T18:53:16.520Z