English

Spoofing-Aware Speaker Verification via Wavelet Prompt Tuning and Multi-Model Ensembles

Audio and Speech Processing 2026-01-27 v1 Sound

Abstract

This paper describes the UZH-CL system submitted to the SASV section of the WildSpoof 2026 challenge. The challenge focuses on the integrated defense against generative spoofing attacks by requiring the simultaneous verification of speaker identity and audio authenticity. We proposed a cascaded Spoofing-Aware Speaker Verification framework that integrates a Wavelet Prompt-Tuned XLSR-AASIST countermeasure with a multi-model ensemble. The ASV component utilizes the ResNet34, ResNet293, and WavLM-ECAPA-TDNN architectures, with Z-score normalization followed by score averaging. Trained on VoxCeleb2 and SpoofCeleb, the system obtained a Macro a-DCF of 0.2017 and a SASV EER of 2.08%. While the system achieved a 0.16% EER in spoof detection on the in-domain data, results on unseen datasets, such as the ASVspoof5, highlight the critical challenge of cross-domain generalization.

Keywords

Cite

@article{arxiv.2601.17557,
  title  = {Spoofing-Aware Speaker Verification via Wavelet Prompt Tuning and Multi-Model Ensembles},
  author = {Aref Farhadipour and Ming Jin and Valeriia Vyshnevetska and Xiyang Li and Elisa Pellegrino and Srikanth Madikeri},
  journal= {arXiv preprint arXiv:2601.17557},
  year   = {2026}
}

Comments

System description of the T03 team in the WildSpoof Challenge at ICASSP 2026

R2 v1 2026-07-01T09:18:43.377Z