English

Baseline Systems for the First Spoofing-Aware Speaker Verification Challenge: Score and Embedding Fusion

Sound 2022-04-22 v1 Audio and Speech Processing

Abstract

Deep learning has brought impressive progress in the study of both automatic speaker verification (ASV) and spoofing countermeasures (CM). Although solutions are mutually dependent, they have typically evolved as standalone sub-systems whereby CM solutions are usually designed for a fixed ASV system. The work reported in this paper aims to gauge the improvements in reliability that can be gained from their closer integration. Results derived using the popular ASVspoof2019 dataset indicate that the equal error rate (EER) of a state-of-the-art ASV system degrades from 1.63% to 23.83% when the evaluation protocol is extended with spoofed trials.%subjected to spoofing attacks. However, even the straightforward integration of ASV and CM systems in the form of score-sum and deep neural network-based fusion strategies reduce the EER to 1.71% and 6.37%, respectively. The new Spoofing-Aware Speaker Verification (SASV) challenge has been formed to encourage greater attention to the integration of ASV and CM systems as well as to provide a means to benchmark different solutions.

Keywords

Cite

@article{arxiv.2204.09976,
  title  = {Baseline Systems for the First Spoofing-Aware Speaker Verification Challenge: Score and Embedding Fusion},
  author = {Hye-jin Shim and Hemlata Tak and Xuechen Liu and Hee-Soo Heo and Jee-weon Jung and Joon Son Chung and Soo-Whan Chung and Ha-Jin Yu and Bong-Jin Lee and Massimiliano Todisco and Héctor Delgado and Kong Aik Lee and Md Sahidullah and Tomi Kinnunen and Nicholas Evans},
  journal= {arXiv preprint arXiv:2204.09976},
  year   = {2022}
}

Comments

8 pages, accepted by Odyssey 2022

R2 v1 2026-06-24T10:54:26.395Z