English

Unsupervised Open-Vocabulary Object Localization in Videos

Computer Vision and Pattern Recognition 2024-06-27 v2

Abstract

In this paper, we show that recent advances in video representation learning and pre-trained vision-language models allow for substantial improvements in self-supervised video object localization. We propose a method that first localizes objects in videos via an object-centric approach with slot attention and then assigns text to the obtained slots. The latter is achieved by an unsupervised way to read localized semantic information from the pre-trained CLIP model. The resulting video object localization is entirely unsupervised apart from the implicit annotation contained in CLIP, and it is effectively the first unsupervised approach that yields good results on regular video benchmarks.

Keywords

Cite

@article{arxiv.2309.09858,
  title  = {Unsupervised Open-Vocabulary Object Localization in Videos},
  author = {Ke Fan and Zechen Bai and Tianjun Xiao and Dominik Zietlow and Max Horn and Zixu Zhao and Carl-Johann Simon-Gabriel and Mike Zheng Shou and Francesco Locatello and Bernt Schiele and Thomas Brox and Zheng Zhang and Yanwei Fu and Tong He},
  journal= {arXiv preprint arXiv:2309.09858},
  year   = {2024}
}

Comments

Accepted by ICCV 2023; Presented on CVPR 2024 Workshop CORR; Project Page:https://github.com/amazon-science/object-centric-vol

R2 v1 2026-06-28T12:24:57.391Z