English

Integrated Spoofing-Robust Automatic Speaker Verification via a Three-Class Formulation and LLR

Audio and Speech Processing 2026-03-19 v2 Sound

Abstract

Spoofing-robust automatic speaker verification (SASV) aims to integrate automatic speaker verification (ASV) and countermeasure (CM). A popular solution is fusion of independent ASV and CM scores. To better modeling SASV, some frameworks integrate ASV and CM within a single network. However, these solutions are typically bi-encoder based, offer limited interpretability, and cannot be readily adapted to new evaluation parameters without retraining. Based on this, we propose a unified end-to-end framework via a three-class formulation that enables log-likelihood ratio (LLR) inference from class logits for a more interpretable decision pipeline. Experiments show comparable performance to existing methods on ASVSpoof5 and better results on SpoofCeleb. The visualization and analysis also prove that the three-class reformulation provides more interpretability.

Keywords

Cite

@article{arxiv.2603.13780,
  title  = {Integrated Spoofing-Robust Automatic Speaker Verification via a Three-Class Formulation and LLR},
  author = {Kai Tan and Lin Zhang and Ruiteng Zhang and Johan Rohdin and Leibny Paola García-Perera and Zexin Cai and Sanjeev Khudanpur and Matthew Wiesner and Nicholas Andrews},
  journal= {arXiv preprint arXiv:2603.13780},
  year   = {2026}
}

Comments

Submitted to Interspeech 2026; put on arxiv based on requirement from Interspeech: "Interspeech no longer enforces an anonymity period for submissions." and "For authors that prefer to upload their paper online, a note indicating that the paper was submitted for review to Interspeech should be included in the posting."

R2 v1 2026-07-01T11:19:46.244Z