English

Online pre-training with long-form videos

Computer Vision and Pattern Recognition 2024-08-29 v1

Abstract

In this study, we investigate the impact of online pre-training with continuous video clips. We will examine three methods for pre-training (masked image modeling, contrastive learning, and knowledge distillation), and assess the performance on downstream action recognition tasks. As a result, online pre-training with contrast learning showed the highest performance in downstream tasks. Our findings suggest that learning from long-form videos can be helpful for action recognition with short videos.

Keywords

Cite

@article{arxiv.2408.15651,
  title  = {Online pre-training with long-form videos},
  author = {Itsuki Kato and Kodai Kamiya and Toru Tamaki},
  journal= {arXiv preprint arXiv:2408.15651},
  year   = {2024}
}

Comments

GCCE2024

R2 v1 2026-06-28T18:26:21.226Z