English

NPU-NTU System for Voice Privacy 2024 Challenge

Audio and Speech Processing 2025-02-05 v2

Abstract

Speaker anonymization is an effective privacy protection solution that conceals the speaker's identity while preserving the linguistic content and paralinguistic information of the original speech. To establish a fair benchmark and facilitate comparison of speaker anonymization systems, the VoicePrivacy Challenge (VPC) was held in 2020 and 2022, with a new edition planned for 2024. In this paper, we describe our proposed speaker anonymization system for VPC 2024. Our system employs a disentangled neural codec architecture and a serial disentanglement strategy to gradually disentangle the global speaker identity and time-variant linguistic content and paralinguistic information. We introduce multiple distillation methods to disentangle linguistic content, speaker identity, and emotion. These methods include semantic distillation, supervised speaker distillation, and frame-level emotion distillation. Based on these distillations, we anonymize the original speaker identity using a weighted sum of a set of candidate speaker identities and a randomly generated speaker identity. Our system achieves the best trade-off of privacy protection and emotion preservation in VPC 2024.

Keywords

Cite

@article{arxiv.2409.04173,
  title  = {NPU-NTU System for Voice Privacy 2024 Challenge},
  author = {Jixun Yao and Nikita Kuzmin and Qing Wang and Pengcheng Guo and Ziqian Ning and Dake Guo and Kong Aik Lee and Eng-Siong Chng and Lei Xie},
  journal= {arXiv preprint arXiv:2409.04173},
  year   = {2025}
}

Comments

System description for VPC 2024

R2 v1 2026-06-28T18:36:20.192Z