English
Related papers

Related papers: InterCap: Joint Markerless 3D Tracking of Humans a…

200 papers

Efficiently detecting human intent to interact with ubiquitous robots is crucial for effective human-robot interaction (HRI) and collaboration. Over the past decade, deep learning has gained traction in this field, with most existing…

Robotics · Computer Science 2025-09-29 Farida Mohsen , Ali Safa

Monocular vertex-level human-scene contact prediction is a fundamental capability for interactive systems such as assistive monitoring, embodied AI, and rehabilitation analysis. In this work, we study this task jointly with single-image 3D…

Computer Vision and Pattern Recognition · Computer Science 2026-03-24 Xiaojian Lin , Yaomin Shen , Junyuan Ma , Yujie Sun , Chengqing Bu , Wenxin Zhang , Zongzheng Zhang , Hao Fei , Lei Jin , Hao Zhao

Recent works in hand-object reconstruction mainly focus on the single-view and dense multi-view settings. On the one hand, single-view methods can leverage learned shape priors to generalise to unseen objects but are prone to inaccuracies…

Computer Vision and Pattern Recognition · Computer Science 2024-05-03 Yik Lung Pang , Changjae Oh , Andrea Cavallaro

In this paper, a marker-based, single-person optical motion capture method (DeepMoCap) is proposed using multiple spatio-temporally aligned infrared-depth sensors and retro-reflective straps and patches (reflectors). DeepMoCap explores…

Computer Vision and Pattern Recognition · Computer Science 2021-10-15 Anargyros Chatzitofis , Dimitrios Zarpalas , Stefanos Kollias , Petros Daras

It is now possible to estimate 3D human pose from monocular images with off-the-shelf 3D pose estimators. However, many practical applications require fine-grained absolute pose information for which multi-view cues and camera calibration…

Computer Vision and Pattern Recognition · Computer Science 2024-05-14 James Tang , Shashwat Suri , Daniel Ajisafe , Bastian Wandt , Helge Rhodin

Articulated objects are prevalent in daily life. Interactable digital twins of such objects have numerous applications in embodied AI and robotics. Unfortunately, current methods to digitize articulated real-world objects require carefully…

Graphics · Computer Science 2025-11-18 Weikun Peng , Jun Lv , Cewu Lu , Manolis Savva

Grasping assistance is essential for restoring autonomy in individuals with motor impairments, particularly in unstructured environments where object categories and user intentions are diverse and unpredictable. We present OVGrasp, a…

Robotics · Computer Science 2025-09-05 Chen Hu , Shan Luo , Letizia Gionfrida

Human performance capture is a highly important computer vision problem with many applications in movie production and virtual/augmented reality. Many previous performance capture approaches either required expensive multi-view setups or…

Computer Vision and Pattern Recognition · Computer Science 2021-11-23 Marc Habermann , Weipeng Xu , Michael Zollhoefer , Gerard Pons-Moll , Christian Theobalt

Marker-based and marker-less optical skeletal motion-capture methods use an outside-in arrangement of cameras placed around a scene, with viewpoints converging on the center. They often create discomfort by possibly needed marker suits, and…

Computer Vision and Pattern Recognition · Computer Science 2016-09-26 Helge Rhodin , Christian Richardt , Dan Casas , Eldar Insafutdinov , Mohammad Shafiei , Hans-Peter Seidel , Bernt Schiele , Christian Theobalt

Reliable markerless motion tracking of people participating in a complex group activity from multiple moving cameras is challenging due to frequent occlusions, strong viewpoint and appearance variations, and asynchronous video streams. To…

Computer Vision and Pattern Recognition · Computer Science 2020-04-21 Minh Vo , Ersin Yumer , Kalyan Sunkavalli , Sunil Hadap , Yaser Sheikh , Srinivasa Narasimhan

We introduce the task of dense captioning in 3D scans from commodity RGB-D sensors. As input, we assume a point cloud of a 3D scene; the expected output is the bounding boxes along with the descriptions for the underlying objects. To…

Computer Vision and Pattern Recognition · Computer Science 2020-12-07 Dave Zhenyu Chen , Ali Gholami , Matthias Nießner , Angel X. Chang

We address the problem of accurate capture and expressive modelling of interactive behaviors happening between two persons in daily scenarios. Different from previous works which either only consider one person or focus on conversational…

Computer Vision and Pattern Recognition · Computer Science 2024-05-28 Yinghao Huang , Leo Ho , Dafei Qin , Mingyi Shi , Taku Komura

In this letter, we present a novel markerless 3D human motion capture (MoCap) system for unstructured, outdoor environments that uses a team of autonomous unmanned aerial vehicles (UAVs) with on-board RGB cameras and computation. Existing…

Computer Vision and Pattern Recognition · Computer Science 2022-04-27 Nitin Saini , Elia Bonetto , Eric Price , Aamir Ahmad , Michael J. Black

We present a novel approach for hand-object action recognition that leverages 2D point tracks as an additional motion cue. While most existing methods rely on RGB appearance, human pose estimation, or their combination, our work…

Computer Vision and Pattern Recognition · Computer Science 2026-01-12 Dennis Holzmann , Sven Wachsmuth

In many visual systems, visual tracking often bases on RGB image sequences, in which some targets are invalid in low-light conditions, and tracking performance is thus affected significantly. Introducing other modalities such as depth and…

Computer Vision and Pattern Recognition · Computer Science 2021-11-12 Chenglong Li , Tianhao Zhu , Lei Liu , Xiaonan Si , Zilin Fan , Sulan Zhai

Reconstructing the motion of objects from videos is a key component for embodied AI and robot manipulation. While diverse approaches to object pose tracking have been studied, they rely heavily on strong external priors, such as depth data…

Computer Vision and Pattern Recognition · Computer Science 2026-05-25 Jisu Shin , Junoh Lee , JunGyu Lee , Inhwan Bae , Dohyeon Lee , Hokyun Im , Youngwoon Lee , Hae-Gon Jeon

Reconstructing compositional 3D representations of scenes, where each object is represented with its own 3D model, is a highly desirable capability in robotics and augmented reality. However, most existing methods rely heavily on strong…

Computer Vision and Pattern Recognition · Computer Science 2024-11-19 Vincent van der Brugge , Marc Pollefeys , Joshua B. Tenenbaum , Ayush Tewari , Krishna Murthy Jatavallabhula

We present a method for inferring diverse 3D models of human-object interactions from images. Reasoning about how humans interact with objects in complex scenes from a single 2D image is a challenging task given ambiguities arising from the…

Computer Vision and Pattern Recognition · Computer Science 2022-09-07 Xi Wang , Gen Li , Yen-Ling Kuo , Muhammed Kocabas , Emre Aksan , Otmar Hilliges

Most prior works in perceiving 3D humans from images reason human in isolation without their surroundings. However, humans are constantly interacting with the surrounding objects, thus calling for models that can reason about not only the…

Computer Vision and Pattern Recognition · Computer Science 2023-11-01 Xianghui Xie , Bharat Lal Bhatnagar , Gerard Pons-Moll

In natural conversation and interaction, our hands often overlap or are in contact with each other. Due to the homogeneous appearance of hands, this makes estimating the 3D pose of interacting hands from images difficult. In this paper we…

Computer Vision and Pattern Recognition · Computer Science 2021-11-30 Zicong Fan , Adrian Spurr , Muhammed Kocabas , Siyu Tang , Michael J. Black , Otmar Hilliges