English

Cross-Modal Bottleneck Fusion For Noise Robust Audio-Visual Speech Recognition

Audio and Speech Processing 2026-02-10 v1

Abstract

Audio-Visual Speech Recognition (AVSR) leverages both acoustic and visual cues to improve speech recognition under noisy conditions. A central question is how to design a fusion mechanism that allows the model to effectively exploit visual information when the audio signal is degraded, while maintaining strong performance on clean speech. We propose CoBRA (Cross-modal Bottleneck for Robust AVSR), a bottleneck-based fusion framework that introduces a compact set of learnable tokens to mediate cross-modal exchange. By regulating information flow through these tokens, the audio stream can reliably access essential visual cues even under adverse or out-of-domain noise. Despite limited training data, our model surpasses comparable baselines and remains competitive with large-scale systems through noise-adaptive fusion, demonstrating both efficiency and robustness. Ablation studies highlight that the depth of fusion is the most critical factor, underscoring its importance in designing robust AVSR systems.

Keywords

Cite

@article{arxiv.2602.08293,
  title  = {Cross-Modal Bottleneck Fusion For Noise Robust Audio-Visual Speech Recognition},
  author = {Seaone Ok and Min Jun Choi and Eungbeom Kim and Seungu Han and Kyogu Lee},
  journal= {arXiv preprint arXiv:2602.08293},
  year   = {2026}
}

Comments

5 pages, 3 figures, ICASSP 2026 Accepted

R2 v1 2026-07-01T10:27:19.086Z