English

Video Swin Transformers for Egocentric Video Understanding @ Ego4D Challenges 2022

Computer Vision and Pattern Recognition 2022-07-26 v1

Abstract

We implemented Video Swin Transformer as a base architecture for the tasks of Point-of-No-Return temporal localization and Object State Change Classification. Our method achieved competitive performance on both challenges.

Keywords

Cite

@article{arxiv.2207.11329,
  title  = {Video Swin Transformers for Egocentric Video Understanding @ Ego4D Challenges 2022},
  author = {Maria Escobar and Laura Daza and Cristina González and Jordi Pont-Tuset and Pablo Arbeláez},
  journal= {arXiv preprint arXiv:2207.11329},
  year   = {2022}
}