English

Joint learning of images and videos with a single Vision Transformer

Computer Vision and Pattern Recognition 2023-08-22 v1

Abstract

In this study, we propose a method for jointly learning of images and videos using a single model. In general, images and videos are often trained by separate models. We propose in this paper a method that takes a batch of images as input to Vision Transformer IV-ViT, and also a set of video frames with temporal aggregation by late fusion. Experimental results on two image datasets and two action recognition datasets are presented.

Keywords

Cite

@article{arxiv.2308.10533,
  title  = {Joint learning of images and videos with a single Vision Transformer},
  author = {Shuki Shimizu and Toru Tamaki},
  journal= {arXiv preprint arXiv:2308.10533},
  year   = {2023}
}

Comments

MVA2023 (18th International Conference on Machine Vision Applications), Hamamatsu, Japan, 23-25 July 2023