English

Self Voice Conversion as an Attack against Neural Audio Watermarking

Sound 2026-03-17 v2 Artificial Intelligence

Abstract

Audio watermarking embeds auxiliary information into speech while maintaining speaker identity, linguistic content, and perceptual quality. Although recent advances in neural and digital signal processing-based watermarking methods have improved imperceptibility and embedding capacity, robustness is still primarily assessed against conventional distortions such as compression, additive noise, and resampling. However, the rise of deep learning-based attacks introduces novel and significant threats to watermark security. In this work, we investigate self voice conversion as a universal, content-preserving attack against audio watermarking systems. Self voice conversion remaps a speaker's voice to the same identity while altering acoustic characteristics through a voice conversion model. We demonstrate that this attack severely degrades the reliability of state-of-the-art watermarking approaches and highlight its implications for the security of modern audio watermarking techniques.

Keywords

Cite

@article{arxiv.2601.20432,
  title  = {Self Voice Conversion as an Attack against Neural Audio Watermarking},
  author = {Yigitcan Özer and Wanying Ge and Zhe Zhang and Xin Wang and Junichi Yamagishi},
  journal= {arXiv preprint arXiv:2601.20432},
  year   = {2026}
}

Comments

7 pages; 2 figures; 2 tables; accepted at IEICE, SP/SLP 2026

R2 v1 2026-07-01T09:23:35.458Z