English
Related papers

Related papers: Learning Coupled Spatial-temporal Attention for Sk…

200 papers

Multi-modality is an important feature of sensor based activity recognition. In this work, we consider two inherent characteristics of human activities, the spatially-temporally varying salience of features and the relations between…

Human-Computer Interaction · Computer Science 2019-05-23 Kaixuan Chen , Lina Yao , Dalin Zhang , Bin Guo , Zhiwen Yu

In this paper, we proposed a effective but extensible residual one-dimensional convolution neural network as base network, based on the this network, we proposed four subnets to explore the features of skeleton sequences from each aspect.…

Computer Vision and Pattern Recognition · Computer Science 2018-08-01 Yangyang Xu , Lei Wang

With the rapid development of digital multimedia, video understanding has become an important field. For action recognition, temporal dimension plays an important role, and this is quite different from image recognition. In order to learn…

Computer Vision and Pattern Recognition · Computer Science 2020-02-11 Qian Liu , Tao Wang , Jie Liu , Yang Guan , Qi Bu , Longfei Yang

Deep learning is ubiquitous across many areas areas of computer vision. It often requires large scale datasets for training before being fine-tuned on small-to-medium scale problems. Activity, or, in other words, action recognition, is one…

Computer Vision and Pattern Recognition · Computer Science 2018-06-26 Yusuf Tas , Piotr Koniusz

Human skeleton joints are popular for action analysis since they can be easily extracted from videos to discard background noises. However, current skeleton representations do not fully benefit from machine learning with CNNs. We propose…

Computer Vision and Pattern Recognition · Computer Science 2018-08-06 Jian Liu , Naveed Akhtar , Ajmal Mian

Deep learning techniques are being used in skeleton based action recognition tasks and outstanding performance has been reported. Compared with RNN based methods which tend to overemphasize temporal information, CNN-based approaches can…

Computer Vision and Pattern Recognition · Computer Science 2017-05-03 Zewei Ding , Pichao Wang , Philip O. Ogunbona , Wanqing Li

Traffic forecasting is one canonical example of spatial-temporal learning task in Intelligent Traffic System. Existing approaches capture spatial dependency with a pre-determined matrix in graph convolution neural operators. However, the…

Machine Learning · Computer Science 2022-06-08 Chen Weikang , Li Yawen , Xue Zhe , Li Ang , Wu Guobin

Event-based moving object detection is a challenging task, where static background and moving object are mixed together. Typically, existing methods mainly align the background events to the same spatial coordinate system via motion…

Computer Vision and Pattern Recognition · Computer Science 2024-03-13 Hanyu Zhou , Zhiwei Shi , Hao Dong , Shihan Peng , Yi Chang , Luxin Yan

Spatiotemporal and motion features are two complementary and crucial information for video action recognition. Recent state-of-the-art methods adopt a 3D CNN stream to learn spatiotemporal features and another flow stream to learn motion…

Computer Vision and Pattern Recognition · Computer Science 2019-08-19 Boyuan Jiang , Mengmeng Wang , Weihao Gan , Wei Wu , Junjie Yan

Vision-based human activity recognition has emerged as one of the essential research areas in video analytics domain. Over the last decade, numerous advanced deep learning algorithms have been introduced to recognize complex human actions…

Computer Vision and Pattern Recognition · Computer Science 2022-08-11 Hayat Ullah , Arslan Munir

Local features at neighboring spatial positions in feature maps have high correlation since their receptive fields are often overlapped. Self-attention usually uses the weighted sum (or other functions) with internal elements of each local…

Computer Vision and Pattern Recognition · Computer Science 2018-08-06 Yang Du , Chunfeng Yuan , Bing Li , Lili Zhao , Yangxi Li , Weiming Hu

Generating video descriptions automatically is a challenging task that involves a complex interplay between spatio-temporal visual features and language models. Given that videos consist of spatial (frame-level) features and their temporal…

Computer Vision and Pattern Recognition · Computer Science 2020-01-20 Anoop Cherian , Jue Wang , Chiori Hori , Tim K. Marks

Medical vision-language pre-training methods mainly leverage the correspondence between paired medical images and radiological reports. Although multi-view spatial images and temporal sequences of image-report pairs are available in…

Artificial Intelligence · Computer Science 2024-05-31 Jinxia Yang , Bing Su , Wayne Xin Zhao , Ji-Rong Wen

As a critical task in video sequence classification within computer vision, Online Action Detection (OAD) has garnered significant attention. The sensitivity of mainstream OAD models to varying video viewpoints often hampers their…

Computer Vision and Pattern Recognition · Computer Science 2025-10-15 Liping Xie , Yang Tan , Shicheng Jing , Huimin Lu , Kanjian Zhang

In skeleton-based human action recognition, temporal pooling is a critical step for capturing spatiotemporal relationship of joint dynamics. Conventional pooling methods overlook the preservation of motion information and treat each frame…

Computer Vision and Pattern Recognition · Computer Science 2024-08-20 Shanaka Ramesh Gunasekara , Wanqing Li , Jack Yang , Philip Ogunbona

This paper proposes a segregated temporal assembly recurrent (STAR) network for weakly-supervised multiple action detection. The model learns from untrimmed videos with only supervision of video-level labels and makes prediction of…

Computer Vision and Pattern Recognition · Computer Science 2018-11-20 Yunlu Xu , Chengwei Zhang , Zhanzhan Cheng , Jianwen Xie , Yi Niu , Shiliang Pu , Fei Wu

Skeleton-based human action recognition has received widespread attention in recent years due to its diverse range of application scenarios. Due to the different sources of human skeletons, skeleton data naturally exhibit heterogeneity. The…

Computer Vision and Pattern Recognition · Computer Science 2025-06-05 Hongsong Wang , Xiaoyan Ma , Jidong Kuang , Jie Gui

In the last years, the computer vision research community has studied on how to model temporal dynamics in videos to employ 3D human action recognition. To that end, two main baseline approaches have been researched: (i) Recurrent Neural…

Computer Vision and Pattern Recognition · Computer Science 2019-09-13 Carlos Caetano , François Brémond , William Robson Schwartz

Skeleton-based action recognition receives increasing attention because the skeleton representations reduce the amount of training data by eliminating visual information irrelevant to actions. To further improve the sample efficiency,…

Computer Vision and Pattern Recognition · Computer Science 2022-09-22 Anqi Zhu , Qiuhong Ke , Mingming Gong , James Bailey

Multi-view action recognition (MVAR) leverages complementary temporal information from different views to improve the learning performance. Obtaining informative view-specific representation plays an essential role in MVAR. Attention has…

Computer Vision and Pattern Recognition · Computer Science 2020-11-30 Yue Bai , Zhiqiang Tao , Lichen Wang , Sheng Li , Yu Yin , Yun Fu