English

Improving Frame-Online Neural Speech Enhancement with Overlapped-Frame Prediction

Sound 2022-07-13 v2 Audio and Speech Processing

Abstract

Frame-online speech enhancement systems in the short-time Fourier transform (STFT) domain usually have an algorithmic latency equal to the window size due to the use of overlap-add in the inverse STFT (iSTFT). This algorithmic latency allows the enhancement models to leverage future contextual information up to a length equal to the window size. However, this information is only partially leveraged by current frame-online systems. To fully exploit it, we propose an overlapped-frame prediction technique for deep learning based frame-online speech enhancement, where at each frame our deep neural network (DNN) predicts the current and several past frames that are necessary for overlap-add, instead of only predicting the current frame. In addition, we propose a loss function to account for the scale difference between predicted and oracle target signals. Experiments on a noisy-reverberant speech enhancement task show the effectiveness of the proposed algorithms.

Keywords

Cite

@article{arxiv.2204.07566,
  title  = {Improving Frame-Online Neural Speech Enhancement with Overlapped-Frame Prediction},
  author = {Zhong-Qiu Wang and Shinji Watanabe},
  journal= {arXiv preprint arXiv:2204.07566},
  year   = {2022}
}

Comments

in IEEE Signal Processing Letters

R2 v1 2026-06-24T10:49:25.062Z