English
Related papers

Related papers: Predicting 4D Hand Trajectory from Monocular Video…

200 papers

Action recognition is essential for egocentric video understanding, allowing automatic and continuous monitoring of Activities of Daily Living (ADLs) without user effort. Existing literature focuses on 3D hand pose input, which requires…

Computer Vision and Pattern Recognition · Computer Science 2024-07-25 Wiktor Mucha , Martin Kampel

Monocular vertex-level human-scene contact prediction is a fundamental capability for interactive systems such as assistive monitoring, embodied AI, and rehabilitation analysis. In this work, we study this task jointly with single-image 3D…

Computer Vision and Pattern Recognition · Computer Science 2026-03-24 Xiaojian Lin , Yaomin Shen , Junyuan Ma , Yujie Sun , Chengqing Bu , Wenxin Zhang , Zongzheng Zhang , Hao Fei , Lei Jin , Hao Zhao

Estimating the 3D pose of a hand is an essential part of human-computer interaction. Estimating 3D pose using depth or multi-view sensors has become easier with recent advances in computer vision, however, regressing pose from a single RGB…

Computer Vision and Pattern Recognition · Computer Science 2018-04-26 Umar Iqbal , Pavlo Molchanov , Thomas Breuel , Juergen Gall , Jan Kautz

We study the problem of estimating the body movements of a camera wearer from egocentric videos. Current methods for ego-body pose estimation rely on temporally dense sensor data, such as IMU measurements from spatially sparse body parts…

Computer Vision and Pattern Recognition · Computer Science 2024-11-07 Seunggeun Chi , Pin-Hao Huang , Enna Sachdeva , Hengbo Ma , Karthik Ramani , Kwonjoon Lee

We present a method for human pose tracking that is based on learning spatiotemporal relationships among joints. Beyond generating the heatmap of a joint in a given frame, our system also learns to predict the offset of the joint from a…

Computer Vision and Pattern Recognition · Computer Science 2019-03-28 Xiao Sun , Chuankang Li , Stephen Lin

Objects manipulated by the hand (i.e., manipulanda) are particularly challenging to reconstruct from Internet videos. Not only does the hand occlude much of the object, but also the object is often only visible in a small number of image…

Computer Vision and Pattern Recognition · Computer Science 2026-01-01 Jane Wu , Georgios Pavlakos , Georgia Gkioxari , Jitendra Malik

In this paper we propose an unsupervised feature extraction method to capture temporal information on monocular videos, where we detect and encode subject of interest in each frame and leverage contrastive self-supervised (CSS) learning to…

Computer Vision and Pattern Recognition · Computer Science 2022-11-28 Sina Honari , Victor Constantin , Helge Rhodin , Mathieu Salzmann , Pascal Fua

Recovering 3D full-body human pose is a challenging problem with many applications. It has been successfully addressed by motion capture systems with body worn markers and multiple cameras. In this paper, we address the more challenging…

Computer Vision and Pattern Recognition · Computer Science 2018-03-12 Xiaowei Zhou , Menglong Zhu , Georgios Pavlakos , Spyridon Leonardos , Kostantinos G. Derpanis , Kostas Daniilidis

Existing multi-person human reconstruction approaches mainly focus on recovering accurate poses or avoiding penetration, but overlook the modeling of close interactions. In this work, we tackle the task of reconstructing closely interactive…

Computer Vision and Pattern Recognition · Computer Science 2024-04-18 Buzhen Huang , Chen Li , Chongyang Xu , Liang Pan , Yangang Wang , Gim Hee Lee

We present the first method to capture the 3D total motion of a target person from a monocular view input. Given an image or a monocular video, our method reconstructs the motion from body, face, and fingers represented by a 3D deformable…

Computer Vision and Pattern Recognition · Computer Science 2018-12-05 Donglai Xiang , Hanbyul Joo , Yaser Sheikh

Hand manipulating objects is an important interaction motion in our daily activities. We faithfully reconstruct this motion with a single RGBD camera by a novel deep reinforcement learning method to leverage physics. Firstly, we propose…

Computer Vision and Pattern Recognition · Computer Science 2024-05-07 Haoyu Hu , Xinyu Yi , Zhe Cao , Jun-Hai Yong , Feng Xu

While 3D hand reconstruction from monocular images has made significant progress, generating accurate and temporally coherent motion estimates from videos remains challenging, particularly during hand-object interactions. In this paper, we…

Computer Vision and Pattern Recognition · Computer Science 2025-08-05 Yufei Zhang , Zijun Cui , Jeffrey O. Kephart , Qiang Ji

We present an approach to estimating camera rotation in crowded, real-world scenes from handheld monocular video. While camera rotation estimation is a well-studied problem, no previous methods exhibit both high accuracy and acceptable…

Computer Vision and Pattern Recognition · Computer Science 2023-09-18 Fabien Delattre , David Dirnfeld , Phat Nguyen , Stephen Scarano , Michael J. Jones , Pedro Miraldo , Erik Learned-Miller

World models have made significant progress in modeling dynamic environments; however, most embodied world models are still restricted to 2D representations, lacking the comprehensive multi-view information essential for embodied spatial…

Computer Vision and Pattern Recognition · Computer Science 2026-05-05 Peiyan Tu , Hanxin Zhu , Jingwen Sun , Shaojie Ren , Cong Wang , Jiayi Luo , Xiaoqian Cheng , Zhibo Chen

The current methods of video-based 3D human pose estimation have achieved significant progress.However, they still face pressing challenges, such as the underutilization of spatiotemporal bodystructure features in transformers and the…

Computer Vision and Pattern Recognition · Computer Science 2025-08-20 Yang Liu , Zhiyong Zhang

Human pose estimation from monocular video is a rapidly advancing field that offers great promise to human movement science and rehabilitation. This potential is tempered by the smaller body of work ensuring the outputs are clinically…

Computer Vision and Pattern Recognition · Computer Science 2022-03-18 R. James Cotton , Emoonah McClerklin , Anthony Cimorelli , Ankit Patel , Tasos Karakostas

Estimating the 3D hand articulation from a single color image is an important problem with applications in Augmented Reality (AR), Virtual Reality (VR), Human-Computer Interaction (HCI), and robotics. Apart from the absence of depth…

Computer Vision and Pattern Recognition · Computer Science 2025-07-18 Christos Pantazopoulos , Spyridon Thermos , Gerasimos Potamianos

We present PAD3R, a method for reconstructing deformable 3D objects from casually captured, unposed monocular videos. Unlike existing approaches, PAD3R handles long video sequences featuring substantial object deformation, large-scale…

Computer Vision and Pattern Recognition · Computer Science 2025-09-30 Ting-Hsuan Liao , Haowen Liu , Yiran Xu , Songwei Ge , Gengshan Yang , Jia-Bin Huang

We tackle the challenging task of estimating global 3D joint locations for both hands via only monocular RGB input images. We propose a novel multi-stage convolutional neural network based pipeline that accurately segments and locates the…

Computer Vision and Pattern Recognition · Computer Science 2020-08-26 Fanqing Lin , Connor Wilhelm , Tony Martinez

We introduce ShapeGaussian, a high-fidelity, template-free method for 4D human reconstruction from casual monocular videos. Generic reconstruction methods lacking robust vision priors, such as 4DGS, struggle to capture high-deformation…

Computer Vision and Pattern Recognition · Computer Science 2026-02-06 Zhenxiao Liang , Ning Zhang , Youbao Tang , Ruei-Sung Lin , Qixing Huang , Peng Chang , Jing Xiao