English

Masked Feature Modelling: Feature Masking for the Unsupervised Pre-training of a Graph Attention Network Block for Bottom-up Video Event Recognition

Computer Vision and Pattern Recognition 2023-08-28 v2 Machine Learning Multimedia

Abstract

In this paper, we introduce Masked Feature Modelling (MFM), a novel approach for the unsupervised pre-training of a Graph Attention Network (GAT) block. MFM utilizes a pretrained Visual Tokenizer to reconstruct masked features of objects within a video, leveraging the MiniKinetics dataset. We then incorporate the pre-trained GAT block into a state-of-the-art bottom-up supervised video-event recognition architecture, ViGAT, to improve the model's starting point and overall accuracy. Experimental evaluations on the YLI-MED dataset demonstrate the effectiveness of MFM in improving event recognition performance.

Keywords

Cite

@article{arxiv.2308.12673,
  title  = {Masked Feature Modelling: Feature Masking for the Unsupervised Pre-training of a Graph Attention Network Block for Bottom-up Video Event Recognition},
  author = {Dimitrios Daskalakis and Nikolaos Gkalelis and Vasileios Mezaris},
  journal= {arXiv preprint arXiv:2308.12673},
  year   = {2023}
}

Comments

8 pages