English
Related papers

Related papers: Unified Keypoint-based Action Recognition Framewor…

200 papers

Point-Level temporal action localization (PTAL) aims to localize actions in untrimmed videos with only one timestamp annotation for each action instance. Existing methods adopt the frame-level prediction paradigm to learn from the sparse…

Computer Vision and Pattern Recognition · Computer Science 2020-12-16 Chen Ju , Peisen Zhao , Ya Zhang , Yanfeng Wang , Qi Tian

One-shot action recognition allows the recognition of human-performed actions with only a single training example. This can influence human-robot-interaction positively by enabling the robot to react to previously unseen behaviour. We…

Computer Vision and Pattern Recognition · Computer Science 2021-03-09 Raphael Memmesheimer , Simon Häring , Nick Theisen , Dietrich Paulus

Self-supervised pretraining methods with masked prediction demonstrate remarkable within-dataset performance in skeleton-based action recognition. However, we show that, unlike contrastive learning approaches, they do not produce…

Computer Vision and Pattern Recognition · Computer Science 2025-11-11 Soroush Mehraban , Mohammad Javad Rajabi , Andrea Iaboni , Babak Taati

Skeleton data, which consists of only the 2D/3D coordinates of the human joints, has been widely studied for human action recognition. Existing methods take the semantics as prior knowledge to group human joints and draw correlations…

Computer Vision and Pattern Recognition · Computer Science 2021-03-23 Lei Shi , Yifan Zhang , Jian Cheng , Hanqing Lu

Existing state-of-the-art 3D point cloud understanding methods merely perform well in a fully supervised manner. To the best of our knowledge, there exists no unified framework that simultaneously solves the downstream high-level…

Computer Vision and Pattern Recognition · Computer Science 2023-12-05 Kangcheng Liu

Human skeleton, as a compact representation of human action, has received increasing attention in recent years. Many skeleton-based action recognition methods adopt graph convolutional networks (GCN) to extract features on top of human…

Computer Vision and Pattern Recognition · Computer Science 2022-04-05 Haodong Duan , Yue Zhao , Kai Chen , Dahua Lin , Bo Dai

The recent multi-modality models have achieved great performance in many vision tasks because the extracted features contain the multi-modality knowledge. However, most of the current registration descriptors have only concentrated on local…

Robotics · Computer Science 2023-02-13 Mingzhi Yuan , Xiaoshui Huang , Kexue Fu , Zhihao Li , Manning Wang

The paper presents a simple and effective learning-based method for computing a discriminative 3D point cloud descriptor for place recognition purposes. Recent state-of-the-art methods have relatively complex architectures such as…

Computer Vision and Pattern Recognition · Computer Science 2022-04-11 Jacek Komorowski

In this study, we present an analysis of model-based ensemble learning for 3D point-cloud object classification and detection. An ensemble of multiple model instances is known to outperform a single model instance, but there is little study…

Computer Vision and Pattern Recognition · Computer Science 2019-05-24 Daniel Koguciuk , Łukasz Chechliński , Tarek El-Gaaly

This paper strives for self-supervised learning of a feature space suitable for skeleton-based action recognition. Our proposal is built upon learning invariances to input skeleton representations and various skeleton augmentations via a…

Computer Vision and Pattern Recognition · Computer Science 2021-08-10 Fida Mohammad Thoker , Hazel Doughty , Cees G. M. Snoek

Rapid progress and superior performance have been achieved for skeleton-based action recognition recently. In this article, we investigate this problem under a cross-dataset setting, which is a new, pragmatic, and challenging task in…

Computer Vision and Pattern Recognition · Computer Science 2022-07-19 Yansong Tang , Xingyu Liu , Xumin Yu , Danyang Zhang , Jiwen Lu , Jie Zhou

Point clouds are a widely available and canonical data modality which convey the 3D geometry of a scene. Despite significant progress in classification and segmentation from point clouds, policy learning from such a modality remains…

Robotics · Computer Science 2022-11-17 Daniel Seita , Yufei Wang , Sarthak J. Shetty , Edward Yao Li , Zackory Erickson , David Held

Test-Time Training (TTT) has emerged as a promising solution to address distribution shifts in 3D point cloud classification. However, existing methods often rely on computationally expensive backpropagation during adaptation, limiting…

Computer Vision and Pattern Recognition · Computer Science 2025-05-27 Ali Bahri , Moslem Yazdanpanah , Sahar Dastani , Mehrdad Noori , Gustavo Adolfo Vargas Hakim , David Osowiechi , Farzad Beizaee , Ismail Ben Ayed , Christian Desrosiers

Current state-of-the-art methods for skeleton-based action recognition are supervised and rely on labels. The reliance is limiting the performance due to the challenges involved in annotation and mislabeled data. Unsupervised methods have…

Computer Vision and Pattern Recognition · Computer Science 2020-12-09 Jingyuan Li , Eli Shlizerman

Deep Learning architectures, albeit successful in most computer vision tasks, were designed for data with an underlying Euclidean structure, which is not usually fulfilled since pre-processed data may lie on a non-linear space. In this…

Computer Vision and Pattern Recognition · Computer Science 2020-11-25 Racha Friji , Hassen Drira , Faten Chaieb , Sebastian Kurtek , Hamza Kchok

We propose a novel skeleton-based representation for 3D action recognition in videos using Deep Convolutional Neural Networks (D-CNNs). Two key issues have been addressed: First, how to construct a robust representation that easily captures…

Computer Vision and Pattern Recognition · Computer Science 2018-07-19 Huy Hieu Pham , Louahdi Khoudour , Alain Crouzil , Pablo Zegers , Sergio A. Velastin

We propose a unified point cloud video self-supervised learning framework for object-centric and scene-centric data. Previous methods commonly conduct representation learning at the clip or frame level and cannot well capture fine-grained…

Computer Vision and Pattern Recognition · Computer Science 2023-08-21 Xiaoxiao Sheng , Zhiqiang Shen , Gang Xiao , Longguang Wang , Yulan Guo , Hehe Fan

We present a deep learning-based multitask framework for joint 3D human pose estimation and action recognition from RGB video sequences. Our approach proceeds along two stages. In the first, we run a real-time 2D pose detector to determine…

Computer Vision and Pattern Recognition · Computer Science 2019-07-17 Huy Hieu Pham , Houssam Salmane , Louahdi Khoudour , Alain Crouzil , Pablo Zegers , Sergio A Velastin

This paper uses clustering algorithms to introduce a shape framework for deformable objects. Until now, the shape detection of the deformable objects has faced several challenges: 1) unable to form a unified framework for multiple shapes;…

Robotics · Computer Science 2023-12-19 Fangqing Chen

Skeleton based action recognition distinguishes human actions using the trajectories of skeleton joints, which provide a very good representation for describing actions. Considering that recurrent neural networks (RNNs) with Long Short-Term…

Computer Vision and Pattern Recognition · Computer Science 2016-03-28 Wentao Zhu , Cuiling Lan , Junliang Xing , Wenjun Zeng , Yanghao Li , Li Shen , Xiaohui Xie