English

Video Object Segmentation via SAM 2: The 4th Solution for LSVOS Challenge VOS Track

Computer Vision and Pattern Recognition 2024-08-27 v2

Abstract

Video Object Segmentation (VOS) task aims to segmenting a particular object instance throughout the entire video sequence given only the object mask of the first frame. Recently, Segment Anything Model 2 (SAM 2) is proposed, which is a foundation model towards solving promptable visual segmentation in images and videos. SAM 2 builds a data engine, which improves model and data via user interaction, to collect the largest video segmentation dataset to date. SAM 2 is a simple transformer architecture with streaming memory for real-time video processing, which trained on the date provides strong performance across a wide range of tasks. In this work, we evaluate the zero-shot performance of SAM 2 on the more challenging VOS datasets MOSE and LVOS. Without fine-tuning on the training set, SAM 2 achieved 75.79 J&F on the test set and ranked 4th place for 6th LSVOS Challenge VOS Track.

Keywords

Cite

@article{arxiv.2408.10125,
  title  = {Video Object Segmentation via SAM 2: The 4th Solution for LSVOS Challenge VOS Track},
  author = {Feiyu Pan and Hao Fang and Runmin Cong and Wei Zhang and Xiankai Lu},
  journal= {arXiv preprint arXiv:2408.10125},
  year   = {2024}
}

Comments

arXiv admin note: substantial text overlap with arXiv:2408.00714

R2 v1 2026-06-28T18:17:00.088Z