English

SAM2Matting: Generalized Image and Video Matting

Computer Vision and Pattern Recognition 2026-06-25 v1

Abstract

Despite impressive advances in image matting, video matting remains challenging due to the inherent gap between high-level tracking, which requires frame-wise understanding, and low-level matting, which focuses on extremely fine-grained details. Existing methods attempt this with expensive and narrowly-scoped video matting datasets, which may limit out-of-domain generalization and compromise tracking robustness. We rethink the paradigm with SAM2Matting, a tracker-to-matting framework that advances VOS trackers to high-fidelity video matting. Specifically, it decouples the task by enhancing a foundational tracker (e.g., SAM2, SAM3) with a region-proposal bridge and dedicated matting heads, enabling the uncompromised tracker to handle temporal consistency while the matting components resolve fine-grained details. Notably, despite being trained only on images, SAM2Matting establishes new state-of-the-art performance on video matting, supports diverse prompt types, maintains strong temporal consistency, and demonstrates robust generalization across both human-centric and in-the-wild scenarios.

Cite

@article{arxiv.2606.27339,
  title  = {SAM2Matting: Generalized Image and Video Matting},
  author = {Ruiqi Shen and Guangquan Jie and Chang Liu and Henghui Ding},
  journal= {arXiv preprint arXiv:2606.27339},
  year   = {2026}
}

Comments

ECCV 2026. Extended version. Project Page: https://henghuiding.com/SAM2Matting/