English

Efficient Audio Enhancement with a Differentiable Psychoacoustic Loss

Audio and Speech Processing 2026-08-03 v1

Abstract

Audio enhancement consists of improving the perceived quality of audio signals. Initially, with the aim of addressing bandwidth extension, this work proposes AEROMambaPAEROMamba_{P}, an efficient variant of the AERO super-resolution architecture where attention and LSTM layers are replaced by the Mamba state-space model, and which incorporates a newly developed differentiable perceptual loss derived from the Perceptual Audio Quality Measure (PAQM). During training, the architecture requires approximately 2-4x less GPU memory than the baseline; during inference, it achieves a 14x speedup while using only one-fifth of the GPU memory. When upsampling both a piano dataset and MUSDB18 from 11.025 kHz to 44.1 kHz, subjective listening tests show that AEROMambaPAEROMamba_{P} outperforms AERO by 15% in perceived quality scores. Next, to handle the enhancement of audio signals that have been highly compressed by lossy coding, it is further proposed AEROMambaPSˉAEROMamba_{P\={S}}, which applies the same framework but replaces STFT reconstruction losses with the PAQM loss, specifically to enhance MP3 encoded audio at 32 kbps. In listening evaluations, AEROMambaPSˉAEROMamba_{P\={S}} achieves 52% higher quality rating than AEROMambaPAEROMamba_{P} when restoring compressed audio. These results demonstrate that PAQM-driven training coupled with lightweight state-space modeling yields high perceptual quality and computational efficiency in both band-limited and compressed audio scenarios.

Keywords

Cite

@article{arxiv.2608.02918,
  title  = {Efficient Audio Enhancement with a Differentiable Psychoacoustic Loss},
  author = {Wallace Abreu and Bernardo V. Miranda and Luiz W. P. Biscainho},
  journal= {arXiv preprint arXiv:2608.02918},
  year   = {2026}
}

Comments

Author's Accepted Manuscript (AAM). Accepted for publication in the Journal of the Audio Engineering Society (JAES) in June/2026