English

Learnable Adaptive Time-Frequency Representation via Differentiable Short-Time Fourier Transform

Sound 2025-06-27 v1 Machine Learning Audio and Speech Processing Signal Processing

Abstract

The short-time Fourier transform (STFT) is widely used for analyzing non-stationary signals. However, its performance is highly sensitive to its parameters, and manual or heuristic tuning often yields suboptimal results. To overcome this limitation, we propose a unified differentiable formulation of the STFT that enables gradient-based optimization of its parameters. This approach addresses the limitations of traditional STFT parameter tuning methods, which often rely on computationally intensive discrete searches. It enables fine-tuning of the time-frequency representation (TFR) based on any desired criterion. Moreover, our approach integrates seamlessly with neural networks, allowing joint optimization of the STFT parameters and network weights. The efficacy of the proposed differentiable STFT in enhancing TFRs and improving performance in downstream tasks is demonstrated through experiments on both simulated and real-world data.

Keywords

Cite

@article{arxiv.2506.21440,
  title  = {Learnable Adaptive Time-Frequency Representation via Differentiable Short-Time Fourier Transform},
  author = {Maxime Leiber and Yosra Marnissi and Axel Barrau and Sylvain Meignen and Laurent Massoulié},
  journal= {arXiv preprint arXiv:2506.21440},
  year   = {2025}
}

Comments

DSTFT, STFT, spectrogram, time-frequency, IEEE Transactions on Signal Processing, 10 pages

R2 v1 2026-07-01T03:34:49.529Z