English

Representation learning from videos in-the-wild: An object-centric approach

Computer Vision and Pattern Recognition 2021-02-10 v2

Abstract

We propose a method to learn image representations from uncurated videos. We combine a supervised loss from off-the-shelf object detectors and self-supervised losses which naturally arise from the video-shot-frame-object hierarchy present in each video. We report competitive results on 19 transfer learning tasks of the Visual Task Adaptation Benchmark (VTAB), and on 8 out-of-distribution-generalization tasks, and discuss the benefits and shortcomings of the proposed approach. In particular, it improves over the baseline on all 18/19 few-shot learning tasks and 8/8 out-of-distribution generalization tasks. Finally, we perform several ablation studies and analyze the impact of the pretrained object detector on the performance across this suite of tasks.

Keywords

Cite

@article{arxiv.2010.02808,
  title  = {Representation learning from videos in-the-wild: An object-centric approach},
  author = {Rob Romijnders and Aravindh Mahendran and Michael Tschannen and Josip Djolonga and Marvin Ritter and Neil Houlsby and Mario Lucic},
  journal= {arXiv preprint arXiv:2010.02808},
  year   = {2021}
}

Comments

Published at WACV 2021

R2 v1 2026-06-23T19:05:32.182Z