English

Unfolded Recursive Expectation-Maximization Neural Network For Speaker Tracking

Audio and Speech Processing 2026-07-29 v1 Signal Processing

Abstract

We propose a deep unfolded REM network for robust tracking of a single moving speaker in mild reverberant environments. Unlike classical REM algorithms, which rely on fixed-step-size decay schedules, the proposed architecture learns an adaptive update policy by unfolding the iterative procedure into differentiable layers. We introduce a Step Size Network that leverages FiLM and PE to dynamically adjust the recursion weights based on temporal context and convergence state. Experimental results for tracking a single speaker under reverberant conditions demonstrate that the proposed unfolded network outperforms the classical CREM baseline, which employs a spatial grid search to map the estimated centroids to physical positions. In the single-speaker tracking task, the proposed method achieves a lower RMSE than the CREM baseline, highlighting its potential for dynamic acoustic scenarios.

Cite

@article{arxiv.2607.26575,
  title  = {Unfolded Recursive Expectation-Maximization Neural Network For Speaker Tracking},
  author = {Rina Veler and Sharon Gannot},
  journal= {arXiv preprint arXiv:2607.26575},
  year   = {2026}
}

Comments

proceedings of IWAENC 2026