Are Modern Speech Enhancement Systems Vulnerable to Adversarial Attacks?
Abstract
Machine learning approaches for speech enhancement are becoming increasingly expressive, enabling ever more powerful modifications of input signals. In this paper, we demonstrate that this expressiveness introduces a vulnerability: advanced speech enhancement models can be susceptible to adversarial attacks. Specifically, we show that adversarial noise, carefully crafted and psychoacoustically masked by the original input, can be injected such that the enhanced speech output conveys an entirely different semantic meaning. We experimentally verify that contemporary predictive speech enhancement models can indeed be manipulated in this way. Furthermore, we highlight that diffusion models with stochastic samplers exhibit inherent robustness to such adversarial attacks by design.
Cite
@article{arxiv.2509.21087,
title = {Are Modern Speech Enhancement Systems Vulnerable to Adversarial Attacks?},
author = {Rostislav Makarov and Lea Schönherr and Timo Gerkmann},
journal= {arXiv preprint arXiv:2509.21087},
year = {2026}
}
Comments
Copyright 2026 IEEE. Personal use of this material is permitted. Permission from IEEE must be obtained for all other uses, in any current or future media, including reprinting/republishing this material for advertising or promotional purposes, creating new collective works, for resale or redistribution to servers or lists, or reuse of any copyrighted component of this work in other works