This technical report presents our approach "Knights" to solve the action recognition task on a small subset of Kinetics-400 i.e. Kinetics400ViPriors without using any extra-data. Our approach has 3 main components: state-of-the-art Temporal Contrastive self-supervised pretraining, video transformer models, and optical flow modality. Along with the use of standard test-time augmentation, our proposed solution achieves 73% on Kinetics400ViPriors test set, which is the best among all of the other entries Visual Inductive Priors for Data-Efficient Computer Vision's Action Recognition Challenge, ICCV 2021.
@article{arxiv.2110.07758,
title = {"Knights": First Place Submission for VIPriors21 Action Recognition Challenge at ICCV 2021},
author = {Ishan Dave and Naman Biyani and Brandon Clark and Rohit Gupta and Yogesh Rawat and Mubarak Shah},
journal= {arXiv preprint arXiv:2110.07758},
year = {2021}
}
Comments
Challenge results are available at https://vipriors.github.io/challenges/#action-recognition