English

DeepActsNet: Spatial and Motion features from Face, Hands, and Body Combined with Convolutional and Graph Networks for Improved Action Recognition

Computer Vision and Pattern Recognition 2021-06-07 v3

Abstract

Existing action recognition methods mainly focus on joint and bone information in human body skeleton data due to its robustness to complex backgrounds and dynamic characteristics of the environments. In this paper, we combine body skeleton data with spatial and motion features from face and two hands, and present "Deep Action Stamps (DeepActs)", a novel data representation to encode actions from video sequences. We also present "DeepActsNet", a deep learning based ensemble model which learns convolutional and structural features from Deep Action Stamps for highly accurate action recognition. Experiments on three challenging action recognition datasets (NTU60, NTU120, and SYSU) show that the proposed model trained using Deep Action Stamps produce considerable improvements in the action recognition accuracy with less computational cost compared to the state-of-the-art methods.

Keywords

Cite

@article{arxiv.2009.09818,
  title  = {DeepActsNet: Spatial and Motion features from Face, Hands, and Body Combined with Convolutional and Graph Networks for Improved Action Recognition},
  author = {Umar Asif and Deval Mehta and Stefan von Cavallar and Jianbin Tang and Stefan Harrer},
  journal= {arXiv preprint arXiv:2009.09818},
  year   = {2021}
}
R2 v1 2026-06-23T18:41:16.256Z