English

LSVOS Challenge 3rd Place Report: SAM2 and Cutie based VOS

Computer Vision and Pattern Recognition 2024-08-22 v2 Information Retrieval

Abstract

Video Object Segmentation (VOS) presents several challenges, including object occlusion and fragmentation, the dis-appearance and re-appearance of objects, and tracking specific objects within crowded scenes. In this work, we combine the strengths of the state-of-the-art (SOTA) models SAM2 and Cutie to address these challenges. Additionally, we explore the impact of various hyperparameters on video instance segmentation performance. Our approach achieves a J\&F score of 0.7952 in the testing phase of LSVOS challenge VOS track, ranking third overall.

Keywords

Cite

@article{arxiv.2408.10469,
  title  = {LSVOS Challenge 3rd Place Report: SAM2 and Cutie based VOS},
  author = {Xinyu Liu and Jing Zhang and Kexin Zhang and Xu Liu and Lingling Li},
  journal= {arXiv preprint arXiv:2408.10469},
  year   = {2024}
}

Comments

arXiv admin note: text overlap with arXiv:2406.03668

R2 v1 2026-06-28T18:17:33.710Z