English

Single and Multi-Speaker Cloned Voice Detection: From Perceptual to Learned Features

Sound 2023-09-28 v2 Computation and Language Audio and Speech Processing

Abstract

Synthetic-voice cloning technologies have seen significant advances in recent years, giving rise to a range of potential harms. From small- and large-scale financial fraud to disinformation campaigns, the need for reliable methods to differentiate real and synthesized voices is imperative. We describe three techniques for differentiating a real from a cloned voice designed to impersonate a specific person. These three approaches differ in their feature extraction stage with low-dimensional perceptual features offering high interpretability but lower accuracy, to generic spectral features, and end-to-end learned features offering less interpretability but higher accuracy. We show the efficacy of these approaches when trained on a single speaker's voice and when trained on multiple voices. The learned features consistently yield an equal error rate between 0% and 4%, and are reasonably robust to adversarial laundering.

Keywords

Cite

@article{arxiv.2307.07683,
  title  = {Single and Multi-Speaker Cloned Voice Detection: From Perceptual to Learned Features},
  author = {Sarah Barrington and Romit Barua and Gautham Koorma and Hany Farid},
  journal= {arXiv preprint arXiv:2307.07683},
  year   = {2023}
}

Comments

S. Barrington, R. Barua, G. Koorma, and Hany Farid. Single and Multi-Speaker Cloned Voice Detection: From Perceptual to Learned Features. Workshop on Image Forensics and Security, Nuremberg, Germany, 2023