English

UW-VOS: A Large-Scale Dataset for Underwater Video Object Segmentation

Computer Vision and Pattern Recognition 2026-03-26 v1

Abstract

Underwater Video Object Segmentation (VOS) is essential for marine exploration, yet open-air methods suffer significant degradation due to color distortion, low contrast, and prevalent camouflage. A primary hurdle is the lack of high-quality training data. To bridge this gap, we introduce UW-VOS\textbf{UW-VOS}, the first large-scale underwater VOS benchmark comprising 1,431 video sequences across 409 categories with 309,295 mask annotations, constructed via a semi-automatic data engine with rigorous human verification. We further propose SAM-U\textbf{SAM-U}, a parameter-efficient framework that adapts SAM2 to the underwater domain. By inserting lightweight adapters into the image encoder, SAM-U achieves state-of-the-art performance with only \sim2%\% trainable parameters. Extensive experiments reveal that existing methods experience an average 13-point J&F\mathcal{J}\&\mathcal{F} drop on UW-VOS, while SAM-U effectively bridges this domain gap. Detailed attribute-based analysis further identifies small targets, camouflage, and exit-re-entry as critical bottlenecks, providing a roadmap for future research in robust underwater perception.

Keywords

Cite

@article{arxiv.2603.24006,
  title  = {UW-VOS: A Large-Scale Dataset for Underwater Video Object Segmentation},
  author = {Hongshen Zhao and Jingkang Tai and Yuhang Wu and Wenkang Zhang and Xi Lan and Shangyan Wang and Tianyu Zhang and Wankou Yang},
  journal= {arXiv preprint arXiv:2603.24006},
  year   = {2026}
}
R2 v1 2026-07-01T11:36:51.277Z