English

Are Deep Speech Denoising Models Robust to Adversarial Noise?

Sound 2026-03-12 v2 Machine Learning Audio and Speech Processing

Abstract

Deep noise suppression (DNS) models enjoy widespread use throughout a variety of high-stakes speech applications. However, we show that four recent DNS models can each be reduced to outputting unintelligible gibberish through the addition of psychoacoustically hidden adversarial noise, even in low-background-noise and simulated over-the-air settings. For three of the models, a small transcription study with audio and multimedia experts confirms unintelligibility of the attacked audio; simultaneously, an ABX study shows that the adversarial noise is generally imperceptible, with some variance between participants and samples. While we also establish several negative results around targeted attacks and model transfer, our results nevertheless highlight the need for practical countermeasures before open-source DNS systems can be used in safety-critical applications.

Keywords

Cite

@article{arxiv.2503.11627,
  title  = {Are Deep Speech Denoising Models Robust to Adversarial Noise?},
  author = {Will Schwarzer and Neel Chaudhari and Philip S. Thomas and Andrea Fanelli and Xiaoyu Liu},
  journal= {arXiv preprint arXiv:2503.11627},
  year   = {2026}
}

Comments

22 pages, 14 figures. Related conference version accepted to ICLR 2026: see https://openreview.net/forum?id=WtH2JxKJKf

R2 v1 2026-06-28T22:20:57.516Z