English

Reference-free Adversarial Sex Obfuscation in Speech

Audio and Speech Processing 2025-08-05 v1 Sound

Abstract

Sex conversion in speech involves privacy risks from data collection and often leaves residual sex-specific cues in outputs, even when target speaker references are unavailable. We introduce RASO for Reference-free Adversarial Sex Obfuscation. Innovations include a sex-conditional adversarial learning framework to disentangle linguistic content from sex-related acoustic markers and explicit regularisation to align fundamental frequency distributions and formant trajectories with sex-neutral characteristics learned from sex-balanced training data. RASO preserves linguistic content and, even when assessed under a semi-informed attack model, it significantly outperforms a competing approach to sex obfuscation.

Cite

@article{arxiv.2508.02295,
  title  = {Reference-free Adversarial Sex Obfuscation in Speech},
  author = {Yangyang Qu and Michele Panariello and Massimiliano Todisco and Nicholas Evans},
  journal= {arXiv preprint arXiv:2508.02295},
  year   = {2025}
}
R2 v1 2026-07-01T04:33:06.146Z