English

Data Quality as Predictor of Voice Anti-Spoofing Generalization

Audio and Speech Processing 2021-06-23 v2 Computer Vision and Pattern Recognition Machine Learning Sound

Abstract

Voice anti-spoofing aims at classifying a given utterance either as a bonafide human sample, or a spoofing attack (e.g. synthetic or replayed sample). Many anti-spoofing methods have been proposed but most of them fail to generalize across domains (corpora) -- and we do not know \emph{why}. We outline a novel interpretative framework for gauging the impact of data quality upon anti-spoofing performance. Our within- and between-domain experiments pool data from seven public corpora and three anti-spoofing methods based on Gaussian mixture and convolutive neural network models. We assess the impacts of long-term spectral information, speaker population (through x-vector speaker embeddings), signal-to-noise ratio, and selected voice quality features.

Keywords

Cite

@article{arxiv.2103.14602,
  title  = {Data Quality as Predictor of Voice Anti-Spoofing Generalization},
  author = {Bhusan Chettri and Rosa González Hautamäki and Md Sahidullah and Tomi Kinnunen},
  journal= {arXiv preprint arXiv:2103.14602},
  year   = {2021}
}

Comments

INTERSPEECH 2021

R2 v1 2026-06-24T00:35:43.659Z