English

Deep-Temporal LSTM for Daily Living Action Recognition

Computer Vision and Pattern Recognition 2018-06-18 v2

Abstract

In this paper, we propose to improve the traditional use of RNNs by employing a many to many model for video classification. We analyze the importance of modeling spatial layout and temporal encoding for daily living action recognition. Many RGB methods focus only on short term temporal information obtained from optical flow. Skeleton based methods on the other hand show that modeling long term skeleton evolution improves action recognition accuracy. In this work, we propose a deep-temporal LSTM architecture which extends standard LSTM and allows better encoding of temporal information. In addition, we propose to fuse 3D skeleton geometry with deep static appearance. We validate our approach on public available CAD60, MSRDailyActivity3D and NTU-RGB+D, achieving competitive performance as compared to the state-of-the art.

Keywords

Cite

@article{arxiv.1802.00421,
  title  = {Deep-Temporal LSTM for Daily Living Action Recognition},
  author = {Srijan Das and Michal Koperski and Francois Bremond and Gianpiero Francesca},
  journal= {arXiv preprint arXiv:1802.00421},
  year   = {2018}
}

Comments

Submitted in conference

R2 v1 2026-06-23T00:07:55.947Z