中文

跨越音频与视觉:通过连接预训练模型实现零样态音视频分割

计算机视觉与模式识别 2025-06-10 v1 声音 音频与语音处理

摘要

Audiovisual segmentation(AVS) aims to identify visual regions corresponding to sound sources,playing a vital role in video understanding, surveillance, and human-computer interaction。Traditional AVS methods depend on large-scale pixel-level annotations,which are costly and time-consuming to obtain。To address this,we propose a novel zero-shot AVS framework that eliminates task-specific training by leveraging multiple pretrained models。Our approach integrates audio, vision, and text representations to bridge modality gaps,enabling precise sound source segmentation without AVS-specific annotations。We systematically explore different strategies for connecting pretrained models and evaluate their efficacy across multiple datasets。Experimental results demonstrate that our framework achieves state-of-the-art zero-shot AVS performance,highlighting the effectiveness of multimodal model integration for finegrained audiovisual segmentation。

关键词

引用

@article{arxiv.2506.06537,
  title  = {Bridging Audio and Vision: Zero-Shot Audiovisual Segmentation by Connecting Pretrained Models},
  author = {Seung-jae Lee and Paul Hongsuck Seo},
  journal= {arXiv preprint arXiv:2506.06537},
  year   = {2025}
}

备注

Accepted on INTERSPEECH2025