English

Higher Order Recurrent Space-Time Transformer for Video Action Prediction

Computer Vision and Pattern Recognition 2021-09-22 v3

Abstract

Endowing visual agents with predictive capability is a key step towards video intelligence at scale. The predominant modeling paradigm for this is sequence learning, mostly implemented through LSTMs. Feed-forward Transformer architectures have replaced recurrent model designs in ML applications of language processing and also partly in computer vision. In this paper we investigate on the competitiveness of Transformer-style architectures for video predictive tasks. To do so we propose HORST, a novel higher order recurrent layer design whose core element is a spatial-temporal decomposition of self-attention for video. HORST achieves state of the art competitive performance on Something-Something early action recognition and EPIC-Kitchens action anticipation, showing evidence of predictive capability that we attribute to our recurrent higher order design of self-attention.

Keywords

Cite

@article{arxiv.2104.08665,
  title  = {Higher Order Recurrent Space-Time Transformer for Video Action Prediction},
  author = {Tsung-Ming Tai and Giuseppe Fiameni and Cheng-Kuang Lee and Oswald Lanz},
  journal= {arXiv preprint arXiv:2104.08665},
  year   = {2021}
}
R2 v1 2026-06-24T01:17:03.430Z