English
Related papers

Related papers: World-Coordinate Human Motion Retargeting via SAM …

200 papers

We propose TRAM, a two-stage method to reconstruct a human's global trajectory and motion from in-the-wild videos. TRAM robustifies SLAM to recover the camera motion in the presence of dynamic humans and uses the scene background to derive…

Computer Vision and Pattern Recognition · Computer Science 2024-09-04 Yufu Wang , Ziyun Wang , Lingjie Liu , Kostas Daniilidis

A dominant paradigm for teaching humanoid robots complex skills is to retarget human motions as kinematic references to train reinforcement learning (RL) policies. However, existing retargeting pipelines often struggle with the significant…

We introduce a novel, data-driven approach for reconstructing temporally coherent 3D motion from unstructured and potentially partial observations of non-rigidly deforming shapes. Our goal is to achieve high-fidelity motion reconstructions…

Computer Vision and Pattern Recognition · Computer Science 2025-09-09 Aymen Merrouche , Stefanie Wuhrer , Edmond Boyer

Human motion provides rich priors for training general-purpose humanoid control policies, but raw demonstrations are often incompatible with a robot's kinematics and dynamics, limiting their direct use. We present a two-stage pipeline for…

Robotics · Computer Science 2026-03-13 Hanwen Wang , Qiayuan Liao , Bike Zhang , Kunzhao Ren , Koushil Sreenath , Xiaobin Xiong

Accurate hand and finger tracking from video has significant clinical applications for monitoring activities of daily living and measuring range of motion, yet monocular video approaches for obtaining hand biomechanics remain…

Computer Vision and Pattern Recognition · Computer Science 2026-05-12 R. James Cotton , Pouyan Firouzabadi , Wendy Murray

We introduce D$^3$-Human, a method for reconstructing Dynamic Disentangled Digital Human geometry from monocular videos. Past monocular video human reconstruction primarily focuses on reconstructing undecoupled clothed human bodies or only…

Computer Vision and Pattern Recognition · Computer Science 2025-01-06 Honghu Chen , Bo Peng , Yunfan Tao , Juyong Zhang

Humanoid robot teleoperation allows humans to integrate their cognitive capabilities with the apparatus to perform tasks that need high strength, manoeuvrability and dexterity. This paper presents a framework for teleoperation of humanoid…

Human mesh recovery (HMR) models 3D human body from monocular videos, with recent works extending it to world-coordinate human trajectory and motion reconstruction. However, most existing methods remain offline, relying on future frames or…

Computer Vision and Pattern Recognition · Computer Science 2026-03-19 Yiwen Zhao , Ce Zheng , Yufu Wang , Hsueh-Han Daniel Yang , Liting Wen , Laszlo A. Jeni

Advances in Deep Learning have recently made it possible to recover full 3D meshes of human poses from individual images. However, extension of this notion to videos for recovering temporally coherent poses still remains unexplored. A major…

Computer Vision and Pattern Recognition · Computer Science 2019-07-02 Jian Liu , Naveed Akhtar , Ajmal Mian

We present the first marker-less approach for temporally coherent 3D performance capture of a human with general clothing from monocular video. Our approach reconstructs articulated human skeleton motion as well as medium-scale non-rigid…

Computer Vision and Pattern Recognition · Computer Science 2018-02-26 Weipeng Xu , Avishek Chatterjee , Michael Zollhöfer , Helge Rhodin , Dushyant Mehta , Hans-Peter Seidel , Christian Theobalt

Retargeting human motion to heterogeneous robots is a fundamental challenge in robotics, primarily due to the severe kinematic and dynamic discrepancies between varying embodiments. Existing solutions typically resort to training…

Robotics · Computer Science 2026-05-27 Haoyu Zhang , Shibo Jin , Lusong Li , Jun Li , Liang Lin , Xiaodong He , Zecui Zeng

We present an approach for 3D global human mesh recovery from monocular videos recorded with dynamic cameras. Our approach is robust to severe and long-term occlusions and tracks human bodies even when they go outside the camera's field of…

Computer Vision and Pattern Recognition · Computer Science 2022-03-31 Ye Yuan , Umar Iqbal , Pavlo Molchanov , Kris Kitani , Jan Kautz

Retargeting human motion to robot poses is a practical approach for teleoperating bimanual humanoid robot arms, but existing methods can be suboptimal and slow, often causing undesirable motion or latency. This is due to optimizing to match…

SAM 3D Body (3DB) achieves state-of-the-art accuracy in monocular 3D human mesh recovery, yet its inference latency of several seconds per image precludes real-time application. We present Fast SAM 3D Body, a training-free acceleration…

Computer Vision and Pattern Recognition · Computer Science 2026-03-17 Timing Yang , Sicheng He , Hongyi Jing , Jiawei Yang , Zhijian Liu , Chuhang Zou , Yue Wang

Accurately reconstructing human behavior in close-interaction scenarios is crucial for enabling realistic virtual interactions in augmented reality, precise motion analysis in sports, and natural collaborative behavior in human-robot tasks.…

Computer Vision and Pattern Recognition · Computer Science 2026-04-16 Qi Xia , Peishan Cong , Ziyi Wang , Yujing Sun , Qin Sun , Xinge Zhu , Mao Ye , Ruigang Yang , Yuexin Ma

This paper describes how to obtain accurate 3D body models and texture of arbitrary people from a single, monocular video in which a person is moving. Based on a parametric body model, we present a robust processing pipeline achieving 3D…

Computer Vision and Pattern Recognition · Computer Science 2018-04-17 Thiemo Alldieck , Marcus Magnor , Weipeng Xu , Christian Theobalt , Gerard Pons-Moll

In this paper, we present self-supervised shared latent embedding (S3LE), a data-driven motion retargeting method that enables the generation of natural motions in humanoid robots from motion capture data or RGB videos. While it requires…

Robotics · Computer Science 2021-03-12 Sungjoon Choi , Min Jae Song , Hyemin Ahn , Joohyung Kim

This paper presents a novel framework that enables real-world humanoid robots to maintain stability while performing human-like motion. Current methods train a policy which allows humanoid robots to follow human body using the massive…

Robotics · Computer Science 2025-05-27 Haoyu Zhao , Sixu Lin , Qingwei Ben , Minyue Dai , Hao Fei , Jingbo Wang , Hua Zou , Junting Dong

Reconstructing people, objects, and their interactions in 3D is a long-standing goal for intelligent systems. Often the input is RGB video from a moving camera, making the task ill-posed; depth is ambiguous, humans and objects occlude each…

Computer Vision and Pattern Recognition · Computer Science 2026-05-19 Lixin Xue , Chengwei Zheng , Georgios Paschalidis , Chen Guo , Manuel Kaufmann , Juan Zarate , Dimitrios Tzionas

Temporal 3D human pose estimation from monocular videos is a challenging task in human-centered computer vision due to the depth ambiguity of 2D-to-3D lifting. To improve accuracy and address occlusion issues, inertial sensor has been…

Computer Vision and Pattern Recognition · Computer Science 2024-04-30 Yiming Bao , Xu Zhao , Dahong Qian