English

Track Anything: Segment Anything Meets Videos

Computer Vision and Pattern Recognition 2023-05-01 v2

Abstract

Recently, the Segment Anything Model (SAM) gains lots of attention rapidly due to its impressive segmentation performance on images. Regarding its strong ability on image segmentation and high interactivity with different prompts, we found that it performs poorly on consistent segmentation in videos. Therefore, in this report, we propose Track Anything Model (TAM), which achieves high-performance interactive tracking and segmentation in videos. To be detailed, given a video sequence, only with very little human participation, i.e., several clicks, people can track anything they are interested in, and get satisfactory results in one-pass inference. Without additional training, such an interactive design performs impressively on video object tracking and segmentation. All resources are available on {https://github.com/gaomingqi/Track-Anything}. We hope this work can facilitate related research.

Keywords

Cite

@article{arxiv.2304.11968,
  title  = {Track Anything: Segment Anything Meets Videos},
  author = {Jinyu Yang and Mingqi Gao and Zhe Li and Shang Gao and Fangjing Wang and Feng Zheng},
  journal= {arXiv preprint arXiv:2304.11968},
  year   = {2023}
}

Comments

Tech-report

R2 v1 2026-06-28T10:15:35.713Z