English

Are Modern Speech Enhancement Systems Vulnerable to Adversarial Attacks?

Audio and Speech Processing 2026-05-01 v3 Machine Learning Sound

Abstract

Machine learning approaches for speech enhancement are becoming increasingly expressive, enabling ever more powerful modifications of input signals. In this paper, we demonstrate that this expressiveness introduces a vulnerability: advanced speech enhancement models can be susceptible to adversarial attacks. Specifically, we show that adversarial noise, carefully crafted and psychoacoustically masked by the original input, can be injected such that the enhanced speech output conveys an entirely different semantic meaning. We experimentally verify that contemporary predictive speech enhancement models can indeed be manipulated in this way. Furthermore, we highlight that diffusion models with stochastic samplers exhibit inherent robustness to such adversarial attacks by design.

Keywords

Cite

@article{arxiv.2509.21087,
  title  = {Are Modern Speech Enhancement Systems Vulnerable to Adversarial Attacks?},
  author = {Rostislav Makarov and Lea Schönherr and Timo Gerkmann},
  journal= {arXiv preprint arXiv:2509.21087},
  year   = {2026}
}

Comments

Copyright 2026 IEEE. Personal use of this material is permitted. Permission from IEEE must be obtained for all other uses, in any current or future media, including reprinting/republishing this material for advertising or promotional purposes, creating new collective works, for resale or redistribution to servers or lists, or reuse of any copyrighted component of this work in other works

R2 v1 2026-07-01T05:56:00.907Z