English

ARiSE: Auto-Regressive Multi-Channel Speech Enhancement

Audio and Speech Processing 2025-06-09 v2

Abstract

We propose ARiSE, an auto-regressive algorithm for multi-channel speech enhancement. ARiSE improves existing deep neural network (DNN) based frame-online multi-channel speech enhancement models by introducing auto-regressive connections, where the estimated target speech at previous frames is leveraged as extra input features to help the DNN estimate the target speech at the current frame. The extra input features can be derived from (a) the estimated target speech in previous frames; and (b) a beamformed mixture with the beamformer computed based on the previous estimated target speech. On the other hand, naively training the DNN in an auto-regressive manner is very slow. To deal with this, we propose a parallel training mechanism to speed up the training. Evaluation results in noisy-reverberant conditions show the effectiveness and potential of the proposed algorithms.

Keywords

Cite

@article{arxiv.2505.22051,
  title  = {ARiSE: Auto-Regressive Multi-Channel Speech Enhancement},
  author = {Pengjie Shen and Xueliang Zhang and Zhong-Qiu Wang},
  journal= {arXiv preprint arXiv:2505.22051},
  year   = {2025}
}

Comments

Accepted by Interspeech 2025

R2 v1 2026-07-01T02:45:31.173Z