English

Object-Centric Learning for Real-World Videos by Predicting Temporal Feature Similarities

Computer Vision and Pattern Recognition 2024-03-18 v2 Machine Learning

Abstract

Unsupervised video-based object-centric learning is a promising avenue to learn structured representations from large, unlabeled video collections, but previous approaches have only managed to scale to real-world datasets in restricted domains. Recently, it was shown that the reconstruction of pre-trained self-supervised features leads to object-centric representations on unconstrained real-world image datasets. Building on this approach, we propose a novel way to use such pre-trained features in the form of a temporal feature similarity loss. This loss encodes semantic and temporal correlations between image patches and is a natural way to introduce a motion bias for object discovery. We demonstrate that this loss leads to state-of-the-art performance on the challenging synthetic MOVi datasets. When used in combination with the feature reconstruction loss, our model is the first object-centric video model that scales to unconstrained video datasets such as YouTube-VIS.

Keywords

Cite

@article{arxiv.2306.04829,
  title  = {Object-Centric Learning for Real-World Videos by Predicting Temporal Feature Similarities},
  author = {Andrii Zadaianchuk and Maximilian Seitzer and Georg Martius},
  journal= {arXiv preprint arXiv:2306.04829},
  year   = {2024}
}

Comments

NeurIPS 2023. Website and code available at https://martius-lab.github.io/videosaur

R2 v1 2026-06-28T10:59:27.863Z