English
Related papers

Related papers: A Spatial-Temporal Transformer based Framework For…

200 papers

Human pose estimation is a critical tool across a variety of healthcare applications. Despite significant progress in pose estimation algorithms targeting adults, such developments for infants remain limited. Existing algorithms for infant…

Computer Vision and Pattern Recognition · Computer Science 2025-04-09 Sarosij Bose , Hannah Dela Cruz , Arindam Dutta , Elena Kokkoni , Konstantinos Karydis , Amit K. Roy-Chowdhury

Shoplifting remains a costly issue for the retail sector, but traditional surveillance systems, which are mostly based on human monitoring, are still largely ineffective, with only about 2% of shoplifters being arrested. Existing AI-based…

Computer Vision and Pattern Recognition · Computer Science 2025-04-29 Narges Rashvand , Ghazal Alinezhad Noghre , Armin Danesh Pazho , Babak Rahimi Ardabili , Hamed Tabkhi

We propose a bootstrapping framework to enhance human optical flow and pose. We show that, for videos involving humans in scenes, we can improve both the optical flow and the pose estimation quality of humans by considering the two tasks at…

Computer Vision and Pattern Recognition · Computer Science 2022-10-31 Aritro Roy Arko , James J. Little , Kwang Moo Yi

Human movement is goal-directed and influenced by the spatial layout of the objects in the scene. To plan future human motion, it is crucial to perceive the environment -- imagine how hard it is to navigate a new room with lights off.…

Computer Vision and Pattern Recognition · Computer Science 2020-08-03 Zhe Cao , Hang Gao , Karttikeya Mangalam , Qi-Zhi Cai , Minh Vo , Jitendra Malik

Online test-time adaptation for 3D human pose estimation is used for video streams that differ from training data. Ground truth 2D poses are used for adaptation, but only estimated 2D poses are available in practice. This paper addresses…

Computer Vision and Pattern Recognition · Computer Science 2025-03-17 Qiuxia Lin , Kerui Gu , Linlin Yang , Angela Yao

Predicting 3D human pose from a single monoscopic video can be highly challenging due to factors such as low resolution, motion blur and occlusion, in addition to the fundamental ambiguity in estimating 3D from 2D. Approaches that directly…

Computer Vision and Pattern Recognition · Computer Science 2021-04-26 Tao Jiang , Necati Cihan Camgoz , Richard Bowden

While head-mounted devices are becoming more compact, they provide egocentric views with significant self-occlusions of the device user. Hence, existing methods often fail to accurately estimate complex 3D poses from egocentric views. In…

Computer Vision and Pattern Recognition · Computer Science 2024-05-16 Hiroyasu Akada , Jian Wang , Vladislav Golyanik , Christian Theobalt

Skeleton-based Human Activity Recognition has achieved great interest in recent years as skeleton data has demonstrated being robust to illumination changes, body scales, dynamic camera views, and complex background. In particular,…

Computer Vision and Pattern Recognition · Computer Science 2021-06-23 Chiara Plizzari , Marco Cannici , Matteo Matteucci

Learning from Demonstration (LfD) provides an intuitive and fast approach to program robotic manipulators. Task parameterized representations allow easy adaptation to new scenes and online observations. However, this approach has been…

Robotics · Computer Science 2021-09-10 An T. Le , Meng Guo , Niels van Duijkeren , Leonel Rozo , Robert Krug , Andras G. Kupcsik , Mathias Buerger

Estimating 3D poses from a monocular video is still a challenging task, despite the significant progress that has been made in recent years. Generally, the performance of existing methods drops when the target person is too small/large, or…

Computer Vision and Pattern Recognition · Computer Science 2020-04-27 Yu Cheng , Bo Yang , Bo Wang , Robby T. Tan

In this paper, we tackle the problem of human motion transfer, where we synthesize novel motion video for a target person that imitates the movement from a reference video. It is a video-to-video translation task in which the estimated…

Computer Vision and Pattern Recognition · Computer Science 2020-04-08 Jian Ren , Menglei Chai , Sergey Tulyakov , Chen Fang , Xiaohui Shen , Jianchao Yang

Motion forecasting plays a pivotal role in autonomous driving systems, enabling vehicles to execute collision warnings and rational local-path planning based on predictions of the surrounding vehicles. However, prevalent methods often…

Computer Vision and Pattern Recognition · Computer Science 2024-07-09 Zhanwen Liu , Chao Li , Nan Yang , Yang Wang , Jiaqi Ma , Guangliang Cheng , Xiangmo Zhao

Precise Event Spotting (PES) in sports videos requires frame-level recognition of fine-grained actions from single-camera footage. Existing PES models typically incorporate lightweight temporal modules such as the Gate Shift Module (GSM) or…

Computer Vision and Pattern Recognition · Computer Science 2025-12-05 Hao Xu , Xinyu Wei , Sam Wells , Sunil Aryal

Despite the great progress in 3D human pose estimation from videos, it is still an open problem to take full advantage of a redundant 2D pose sequence to learn representative representations for generating one 3D pose. To this end, we…

Computer Vision and Pattern Recognition · Computer Science 2022-01-12 Wenhao Li , Hong Liu , Runwei Ding , Mengyuan Liu , Pichao Wang , Wenming Yang

Human motion prediction, i.e., forecasting future body poses given observed pose sequence, has typically been tackled with recurrent neural networks (RNNs). However, as evidenced by prior work, the resulted RNN models suffer from prediction…

Computer Vision and Pattern Recognition · Computer Science 2020-07-08 Wei Mao , Miaomiao Liu , Mathieu Salzmann , Hongdong Li

Human pose analysis is presently dominated by deep convolutional networks trained with extensive manual annotations of joint locations and beyond. To avoid the need for expensive labeling, we exploit spatiotemporal relations in training…

Computer Vision and Pattern Recognition · Computer Science 2017-08-08 Ömer Sümer , Tobias Dencker , Björn Ommer

We address human action recognition from multi-modal video data involving articulated pose and RGB frames and propose a two-stream approach. The pose stream is processed with a convolutional model taking as input a 3D tensor holding data…

Computer Vision and Pattern Recognition · Computer Science 2017-08-08 Fabien Baradel , Christian Wolf , Julien Mille

Video-based human pose estimation in crowded scenes is a challenging problem due to occlusion, motion blur, scale variation and viewpoint change, etc. Prior approaches always fail to deal with this problem because of (1) lacking of usage of…

Computer Vision and Pattern Recognition · Computer Science 2020-10-22 Li Yuan , Shuning Chang , Xuecheng Nie , Ziyuan Huang , Yichen Zhou , Yunpeng Chen , Jiashi Feng , Shuicheng Yan

Human motion prediction combines the tasks of trajectory forecasting and human pose prediction. For each of the two tasks, specialized models have been developed. Combining these models for holistic human motion prediction is non-trivial,…

Computer Vision and Pattern Recognition · Computer Science 2026-03-05 Aadya Agrawal , Alexander Schwing

Existing 3D human pose estimation algorithms trained on distortion-free datasets suffer performance drop when applied to new scenarios with a specific camera distortion. In this paper, we propose a simple yet effective model for 3D human…

Computer Vision and Pattern Recognition · Computer Science 2021-12-06 Hanbyel Cho , Yooshin Cho , Jaemyung Yu , Junmo Kim