English

Learning What To Hear: Boosting Sound-Source Association For Robust Audiovisual Instance Segmentation

Audio and Speech Processing 2026-01-30 v2 Artificial Intelligence Multimedia Sound

Abstract

Audiovisual instance segmentation (AVIS) requires accurately localizing and tracking sounding objects throughout video sequences. Existing methods suffer from visual bias stemming from two fundamental issues: uniform additive fusion prevents queries from specializing to different sound sources, while visual-only training objectives allow queries to converge to arbitrary salient objects. We propose Audio-Centric Query Generation using cross-attention, enabling each query to selectively attend to distinct sound sources and carry sound-specific priors into visual decoding. Additionally, we introduce Sound-Aware Ordinal Counting (SAOC) loss that explicitly supervises sounding object numbers through ordinal regression with monotonic consistency constraints, preventing visual-only convergence during training. Experiments on AVISeg benchmark demonstrate consistent improvements: +1.64 mAP, +0.6 HOTA, and +2.06 FSLA, validating that query specialization and explicit counting supervision are crucial for accurate audiovisual instance segmentation.

Keywords

Cite

@article{arxiv.2509.22740,
  title  = {Learning What To Hear: Boosting Sound-Source Association For Robust Audiovisual Instance Segmentation},
  author = {Jinbae Seo and Hyeongjun Kwon and Kwonyoung Kim and Jiyoung Lee and Kwanghoon Sohn},
  journal= {arXiv preprint arXiv:2509.22740},
  year   = {2026}
}

Comments

Accepted to ICASSP 2026

R2 v1 2026-07-01T05:59:33.581Z