English

Generating gender-ambiguous voices for privacy-preserving speech recognition

Sound 2022-07-05 v1 Machine Learning Audio and Speech Processing

Abstract

Our voice encodes a uniquely identifiable pattern which can be used to infer private attributes, such as gender or identity, that an individual might wish not to reveal when using a speech recognition service. To prevent attribute inference attacks alongside speech recognition tasks, we present a generative adversarial network, GenGAN, that synthesises voices that conceal the gender or identity of a speaker. The proposed network includes a generator with a U-Net architecture that learns to fool a discriminator. We condition the generator only on gender information and use an adversarial loss between signal distortion and privacy preservation. We show that GenGAN improves the trade-off between privacy and utility compared to privacy-preserving representation learning methods that consider gender information as a sensitive attribute to protect.

Keywords

Cite

@article{arxiv.2207.01052,
  title  = {Generating gender-ambiguous voices for privacy-preserving speech recognition},
  author = {Dimitrios Stoidis and Andrea Cavallaro},
  journal= {arXiv preprint arXiv:2207.01052},
  year   = {2022}
}

Comments

5 pages, 4 figures, submitted to INTERSPEECH

R2 v1 2026-06-24T12:12:28.152Z