English

Linguistic Bias Mitigation for Spoofing Detection via Gradient Reversal and A Variational Information Bottleneck

Computation and Language 2026-06-30 v1 Machine Learning

Abstract

Rapid advancements in generative speech technology have compromised the reliability of voice biometrics. While current spoofing detectors excel when assessed under in-domain conditions, generalisation to out-of-domain settings is often poor. We show that this can be due to linguistic bias. A reliance on linguistic cues observed in training data can then compromise robustness to cross-data. We propose a linguistic-invariant spoofing detection framework utilizing teacher-student adversarial learning. The linguistic-aware teacher model, pre-trained on linguistic content of an external dataset, guides the student detector via gradient reversal to minimize the linguistic information. To prevent the inadvertent removal of non-linguistic cues, we incorporate a Variational Information Bottleneck to enable suppression of principal cues. Across nine DF Arena datasets, our method achieves up to a 36.2% relative reduction in the EER compare to the baseline.

Keywords

Cite

@article{arxiv.2606.31411,
  title  = {Linguistic Bias Mitigation for Spoofing Detection via Gradient Reversal and A Variational Information Bottleneck},
  author = {Anh-Tuan Dao and Driss Matrouf and Mickael Rouvier and Nicholas Evans},
  journal= {arXiv preprint arXiv:2606.31411},
  year   = {2026}
}