We present a number of systems for the Voice Privacy Challenge, including voice conversion based systems such as the kNN-VC method and the WavLM voice Conversion method, and text-to-speech (TTS) based systems including Whisper-VITS. We found that while voice conversion systems better preserve emotional content, they struggle to conceal speaker identity in semi-white-box attack scenarios; conversely, TTS methods perform better at anonymization and worse at emotion preservation. Finally, we propose a random admixture system which seeks to balance out the strengths and weaknesses of the two category of systems, achieving a strong EER of over 40% while maintaining UAR at a respectable 47%.
@article{arxiv.2409.08913,
title = {HLTCOE JHU Submission to the Voice Privacy Challenge 2024},
author = {Henry Li Xinyuan and Zexin Cai and Ashi Garg and Kevin Duh and Leibny Paola García-Perera and Sanjeev Khudanpur and Nicholas Andrews and Matthew Wiesner},
journal= {arXiv preprint arXiv:2409.08913},
year = {2024}
}
Comments
Submission to the Voice Privacy Challenge 2024. Accepted and presented at the 4th Symposium on Security and Privacy in Speech Communication