English

Real-time Denoising and Dereverberation with Tiny Recurrent U-Net

Sound 2021-06-24 v3 Artificial Intelligence Audio and Speech Processing

Abstract

Modern deep learning-based models have seen outstanding performance improvement with speech enhancement tasks. The number of parameters of state-of-the-art models, however, is often too large to be deployed on devices for real-world applications. To this end, we propose Tiny Recurrent U-Net (TRU-Net), a lightweight online inference model that matches the performance of current state-of-the-art models. The size of the quantized version of TRU-Net is 362 kilobytes, which is small enough to be deployed on edge devices. In addition, we combine the small-sized model with a new masking method called phase-aware β\beta-sigmoid mask, which enables simultaneous denoising and dereverberation. Results of both objective and subjective evaluations have shown that our model can achieve competitive performance with the current state-of-the-art models on benchmark datasets using fewer parameters by orders of magnitude.

Keywords

Cite

@article{arxiv.2102.03207,
  title  = {Real-time Denoising and Dereverberation with Tiny Recurrent U-Net},
  author = {Hyeong-Seok Choi and Sungjin Park and Jie Hwan Lee and Hoon Heo and Dongsuk Jeon and Kyogu Lee},
  journal= {arXiv preprint arXiv:2102.03207},
  year   = {2021}
}

Comments

5 pages, 2 figures, 2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). arXiv admin note: text overlap with arXiv:2006.00687

R2 v1 2026-06-23T22:52:32.105Z