English

An Intervention-Based Framework for Shortcut Diagnosis in Spoofing Countermeasures

Audio and Speech Processing 2026-07-03 v1 Machine Learning

Abstract

While deepfake audio detection systems achieve high performance in controlled benchmarks, their reliability often diminishes in the wild. Prior work shows that dataset-specific artifacts contribute to this gap. Yet, systematic tools to identify which acoustic properties a model exploits as shortcuts remain limited. We propose an intervention-based diagnostic framework, grounded in a directed graphical model, that formally distinguishes confound-driven shortcut dependencies from legitimate domain shift. We operationalise this through controlled acoustic perturbations targeting non-speech structure, spectral content, and signal energy, complemented by corpus-level distributional analysis. Evaluating XLS-R-300M with RawGAT-ST across ASVspoof challenges datasets, we quantify model sensitivity to specific intervention types. Results reveal that non-speech interventions produce the largest performance shifts, confirming non-speech intervals as a dominant shortcut.

Cite

@article{arxiv.2607.03150,
  title  = {An Intervention-Based Framework for Shortcut Diagnosis in Spoofing Countermeasures},
  author = {Santiago Rubio and Pilar Bello and Dayana Ribas and Antonio Miguel and Eduardo Lleida and Alfonso Ortega},
  journal= {arXiv preprint arXiv:2607.03150},
  year   = {2026}
}

Comments

Accepted at Odyssey 2026: The Speaker and Language Recognition Workshop

R2 v1 2026-07-22T20:24:24.766Z