English
Related papers

Related papers: Heatmap Pooling Network for Action Recognition fro…

200 papers

RGB-D salient object detection (SOD) aims to detect the prominent regions by jointly modeling RGB and depth information. Most RGB-D SOD methods apply the same type of backbones and fusion modules to identically learn the multimodality and…

Computer Vision and Pattern Recognition · Computer Science 2023-07-04 Kang Yi , Jing Xu , Xiao Jin , Fu Guo , Yan-Feng Wu

Signals from different modalities each have their own combination algebra which affects their sampling processing. RGB is mostly linear; depth is a geometric signal following the operations of mathematical morphology. If a network obtaining…

Computer Vision and Pattern Recognition · Computer Science 2023-10-12 Rick Groenendijk , Leo Dorst , Theo Gevers

Training robust deep video representations has proven to be much more challenging than learning deep image representations. This is in part due to the enormous size of raw video streams and the high temporal redundancy; the true and…

Computer Vision and Pattern Recognition · Computer Science 2018-04-02 Chao-Yuan Wu , Manzil Zaheer , Hexiang Hu , R. Manmatha , Alexander J. Smola , Philipp Krähenbühl

Robust visual tracking is a challenging computer vision problem, with many real-world applications. Most existing approaches employ hand-crafted appearance features, such as HOG or Color Names. Recently, deep RGB features extracted from…

Computer Vision and Pattern Recognition · Computer Science 2016-12-21 Susanna Gladh , Martin Danelljan , Fahad Shahbaz Khan , Michael Felsberg

Video activity recognition by deep neural networks is impressive for many classes. However, it falls short of human performance, especially for challenging to discriminate activities. Humans differentiate these complex activities by…

Computer Vision and Pattern Recognition · Computer Science 2022-01-12 Joseph Chrol-Cannon , Andrew Gilbert , Ranko Lazic , Adithya Madhusoodanan , Frank Guerin

We investigate the problem of representing an entire video using CNN features for human action recognition. Currently, limited by GPU memory, we have not been able to feed a whole video into CNN/RNNs for end-to-end learning. A common…

Computer Vision and Pattern Recognition · Computer Science 2017-01-31 Zhenzhong Lan , Yi Zhu , Alexander G. Hauptmann

Human Action Recognition (HAR) is a challenging domain in computer vision, involving recognizing complex patterns by analyzing the spatiotemporal dynamics of individuals' movements in videos. These patterns arise in sequential data, such as…

Computer Vision and Pattern Recognition · Computer Science 2025-01-23 Ali K. AlShami , Ryan Rabinowitz , Khang Lam , Yousra Shleibik , Melkamu Mersha , Terrance Boult , Jugal Kalita

This paper presents a 2D skeleton-based action segmentation method with applications in fine-grained human activity recognition. In contrast with state-of-the-art methods which directly take sequences of 3D skeleton coordinates as inputs…

Computer Vision and Pattern Recognition · Computer Science 2024-04-29 Syed Waleed Hyder , Muhammad Usama , Anas Zafar , Muhammad Naufil , Fawad Javed Fateh , Andrey Konin , M. Zeeshan Zia , Quoc-Huy Tran

The demand for accurate on-device pattern recognition in edge applications is intensifying, yet existing approaches struggle to reconcile accuracy with computational constraints. To address this challenge, a resource-aware hierarchical…

Computer Vision and Pattern Recognition · Computer Science 2026-02-03 Boyu Li , Kuangji Zuo , Lincong Li , Yonghui Wu

RGB-thermal salient object detection (RGB-T SOD) aims to locate the common prominent objects of an aligned visible and thermal infrared image pair and accurately segment all the pixels belonging to those objects. It is promising in…

Computer Vision and Pattern Recognition · Computer Science 2022-07-11 Xiurong Jiang , Lin Zhu , Yifan Hou , Hui Tian

Deep hashing approaches, including deep quantization and deep binary hashing, have become a common solution to large-scale image retrieval due to their high computation and storage efficiency. Most existing hashing methods cannot produce…

Computer Vision and Pattern Recognition · Computer Science 2024-01-10 Ziyun Zeng , Jinpeng Wang , Bin Chen , Tao Dai , Shu-Tao Xia , Zhi Wang

Different from RGB videos, depth data in RGB-D videos provide key complementary information for tristimulus visual data which potentially could achieve accuracy improvement for action recognition. However, most of the existing action…

Computer Vision and Pattern Recognition · Computer Science 2018-11-27 Haokui Zhang , Ying Li , Peng Wang , Yu Liu , Chunhua Shen

Human action recognition (HAR) in videos is one of the core tasks of video understanding. Based on video sequences, the goal is to recognize actions performed by humans. While HAR has received much attention in the visible spectrum, action…

Computer Vision and Pattern Recognition · Computer Science 2022-04-20 Soufiane Lamghari , Guillaume-Alexandre Bilodeau , Nicolas Saunier

Presenting high-resolution (HR) human appearance is always critical for the human-centric videos. However, current imagery equipment can hardly capture HR details all the time. Existing super-resolution algorithms barely mitigate the…

Computer Vision and Pattern Recognition · Computer Science 2020-05-28 Guanghan Li , Yaping Zhao , Mengqi Ji , Xiaoyun Yuan , Lu Fang

Skeleton-based human action recognition (HAR) has achieved remarkable progress with graph-based architectures. However, most existing methods remain body-centric, focusing on large-scale motions while neglecting subtle hand articulations…

Computer Vision and Pattern Recognition · Computer Science 2026-01-05 Seungyeon Cho , Tae-kyun Kim

Deep neural networks have recently achieved competitive accuracy for human activity recognition. However, there is room for improvement, especially in modeling long-term temporal importance and determining the activity relevance of…

Computer Vision and Pattern Recognition · Computer Science 2018-08-23 Sibo Song , Ngai-Man Cheung , Vijay Chandrasekhar , Bappaditya Mandal

In this report, our approach to tackling the task of ActivityNet 2018 Kinetics-600 challenge is described in detail. Though spatial-temporal modelling methods, which adopt either such end-to-end framework as I3D \cite{i3d} or two-stage…

Computer Vision and Pattern Recognition · Computer Science 2018-06-28 Dongliang He , Fu Li , Qijie Zhao , Xiang Long , Yi Fu , Shilei Wen

Most action recognition methods base on a) a late aggregation of frame level CNN features using average pooling, max pooling, or RNN, among others, or b) spatio-temporal aggregation via 3D convolutions. The first assume independence among…

Computer Vision and Pattern Recognition · Computer Science 2019-05-30 Swathikiran Sudhakaran , Sergio Escalera , Oswald Lanz

In this work, we present novel temporal encoding methods for action and activity classification by extending the unsupervised rank pooling temporal encoding method in two ways. First, we present "discriminative rank pooling" in which the…

Computer Vision and Pattern Recognition · Computer Science 2017-05-31 Basura Fernando , Stephen Gould

In this work, we study a novel problem which focuses on person identification while performing daily activities. Learning biometric features from RGB videos is challenging due to spatio-temporal complexity and presence of appearance biases…

Computer Vision and Pattern Recognition · Computer Science 2024-03-27 Shehreen Azad , Yogesh Singh Rawat