English

Leveraging Mamba with Full-Face Vision for Audio-Visual Speech Enhancement

Sound 2025-10-01 v2 Audio and Speech Processing

Abstract

Recent Mamba-based models have shown promise in speech enhancement by efficiently modeling long-range temporal dependencies. However, models like Speech Enhancement Mamba (SEMamba) remain limited to single-speaker scenarios and struggle in complex multi-speaker environments such as the cocktail party problem. To overcome this, we introduce AVSEMamba, an audio-visual speech enhancement model that integrates full-face visual cues with a Mamba-based temporal backbone. By leveraging spatiotemporal visual information, AVSEMamba enables more accurate extraction of target speech in challenging conditions. Evaluated on the AVSEC-4 Challenge development and blind test sets, AVSEMamba outperforms other monaural baselines in speech intelligibility (STOI), perceptual quality (PESQ), and non-intrusive quality (UTMOS), and achieves \textbf{1st place} on the monaural leaderboard.

Keywords

Cite

@article{arxiv.2508.13624,
  title  = {Leveraging Mamba with Full-Face Vision for Audio-Visual Speech Enhancement},
  author = {Rong Chao and Wenze Ren and You-Jin Li and Kuo-Hsuan Hung and Sung-Feng Huang and Szu-Wei Fu and Wen-Huang Cheng and Yu Tsao},
  journal= {arXiv preprint arXiv:2508.13624},
  year   = {2025}
}

Comments

Accepted to Interspeech 2025 Workshop

R2 v1 2026-07-01T04:56:19.480Z