English

Is an Object-Centric Video Representation Beneficial for Transfer?

Computer Vision and Pattern Recognition 2022-10-11 v2

Abstract

The objective of this work is to learn an object-centric video representation, with the aim of improving transferability to novel tasks, i.e., tasks different from the pre-training task of action classification. To this end, we introduce a new object-centric video recognition model based on a transformer architecture. The model learns a set of object-centric summary vectors for the video, and uses these vectors to fuse the visual and spatio-temporal trajectory 'modalities' of the video clip. We also introduce a novel trajectory contrast loss to further enhance objectness in these summary vectors. With experiments on four datasets -- SomethingSomething-V2, SomethingElse, Action Genome and EpicKitchens -- we show that the object-centric model outperforms prior video representations (both object-agnostic and object-aware), when: (1) classifying actions on unseen objects and unseen environments; (2) low-shot learning of novel classes; (3) linear probe to other downstream tasks; as well as (4) for standard action classification.

Keywords

Cite

@article{arxiv.2207.10075,
  title  = {Is an Object-Centric Video Representation Beneficial for Transfer?},
  author = {Chuhan Zhang and Ankush Gupta and Andrew Zisserman},
  journal= {arXiv preprint arXiv:2207.10075},
  year   = {2022}
}

Comments

Accepted to ACCV 2022

R2 v1 2026-06-25T01:05:32.196Z