HLTCOE JHU 提交至 2024 年语音隐私挑战赛的系统
音频与语音处理
2024-09-18 v2 机器学习
摘要
我们为语音隐私挑战赛提出了多种系统,包括基于语音转换的系统,如 kNN-VC 方法和 WavLM 语音转换方法,以及基于文本到语音(TTS)的系统,包括 Whisper-VITS。我们发现,虽然语音转换系统能更好地保留情感内容,但在半白盒攻击场景下难以隐藏说话人身份;相反,TTS 方法在匿名化方面表现更好,但在情感保留方面较差。最后,我们提出了一种随机混合系统,旨在平衡这两类系统的优缺点,在保持 UAR 达到可观的 47% 的同时,实现了超过 40% 的强 EER。
关键词
引用
@article{arxiv.2409.08913,
title = {HLTCOE JHU Submission to the Voice Privacy Challenge 2024},
author = {Henry Li Xinyuan and Zexin Cai and Ashi Garg and Kevin Duh and Leibny Paola García-Perera and Sanjeev Khudanpur and Nicholas Andrews and Matthew Wiesner},
journal= {arXiv preprint arXiv:2409.08913},
year = {2024}
}
备注
Submission to the Voice Privacy Challenge 2024. Accepted and presented at the 4th Symposium on Security and Privacy in Speech Communication