English
Related papers

Related papers: MaskHOI: Robust 3D Hand-Object Interaction Estimat…

200 papers

We propose a robust and accurate method for reconstructing 3D hand mesh from monocular images. This is a very challenging problem, as hands are often severely occluded by objects. Previous works often have disregarded 2D hand pose…

Computer Vision and Pattern Recognition · Computer Science 2024-03-14 Shuaibing Wang , Shunli Wang , Dingkang Yang , Mingcheng Li , Ziyun Qian , Liuzhen Su , Lihua Zhang

3D interacting hand reconstruction is essential to facilitate human-machine interaction and human behaviors understanding. Previous works in this field either rely on auxiliary inputs such as depth images or they can only handle a single…

Computer Vision and Pattern Recognition · Computer Science 2021-11-02 Yu Rong , Jingbo Wang , Ziwei Liu , Chen Change Loy

Object pose estimation is a fundamental task in 3D vision with applications in robotics, AR/VR, and scene understanding. We address the challenge of category-level 9-DoF pose estimation (6D pose + 3Dsize) from RGB-D input, without relying…

Computer Vision and Pattern Recognition · Computer Science 2026-03-31 Rachit Agarwal , Abhishek Joshi , Sathish Chalasani , Woo Jin Kim

Occlusion poses a great threat to monocular multi-person 3D human pose estimation due to large variability in terms of the shape, appearance, and position of occluders. While existing methods try to handle occlusion with pose…

Computer Vision and Pattern Recognition · Computer Science 2022-08-02 Qihao Liu , Yi Zhang , Song Bai , Alan Yuille

Automatic estimation of 3D human pose from monocular RGB images is a challenging and unsolved problem in computer vision. In a supervised manner, approaches heavily rely on laborious annotations and present hampered generalization ability…

Computer Vision and Pattern Recognition · Computer Science 2024-07-09 Yuchen Yang , Yu Qiao , Xiao Sun

Masked Autoencoders (MAE) have been popular paradigms for large-scale vision representation pre-training. However, MAE solely reconstructs the low-level RGB signals after the decoder and lacks supervision upon high-level semantics for the…

Computer Vision and Pattern Recognition · Computer Science 2023-03-10 Peng Gao , Renrui Zhang , Rongyao Fang , Ziyi Lin , Hongyang Li , Hongsheng Li , Qiao Yu

Current perception models in autonomous driving heavily rely on large-scale labelled 3D data, which is both costly and time-consuming to annotate. This work proposes a solution to reduce the dependence on labelled 3D training data by…

Computer Vision and Pattern Recognition · Computer Science 2023-10-10 Chen Min , Xinli Xu , Dawei Zhao , Liang Xiao , Yiming Nie , Bin Dai

Rapid progress has been witnessed for human-object interaction (HOI) recognition, but most existing models are confined to single-stage reasoning pipelines. Considering the intrinsic complexity of the task, we introduce a cascade…

Computer Vision and Pattern Recognition · Computer Science 2020-03-26 Tianfei Zhou , Wenguan Wang , Siyuan Qi , Haibin Ling , Jianbing Shen

Current digital human studies focusing on lip-syncing and body movement are no longer sufficient to meet the growing industrial demand, while human video generation techniques that support interacting with real-world environments (e.g.,…

Computer Vision and Pattern Recognition · Computer Science 2025-03-26 Yingying Fan , Quanwei Yang , Kaisiyuan Wang , Hang Zhou , Yingying Li , Haocheng Feng , Errui Ding , Yu Wu , Jingdong Wang

Recently, multi-modal masked autoencoders (MAE) has been introduced in 3D self-supervised learning, offering enhanced feature learning by leveraging both 2D and 3D data to capture richer cross-modal representations. However, these…

Computer Vision and Pattern Recognition · Computer Science 2025-07-29 Zhimin Chen , Xuewei Chen , Xiao Guo , Yingwei Li , Longlong Jing , Liang Yang , Bing Li

Masked Autoencoders (MAEs) have been shown to be effective in pre-training Vision Transformers (ViTs) for natural and medical image analysis problems. By reconstructing missing pixel/voxel information in visible patches, a ViT encoder can…

Computer Vision and Pattern Recognition · Computer Science 2025-11-20 Pengfei Gu , Huimin Li , Yejia Zhang , Chaoli Wang , Danny Z. Chen

Monocular 4D human-object interaction (HOI) reconstruction - recovering a moving human and a manipulated object from a single RGB video - remains challenging due to depth ambiguity and frequent occlusions. Existing methods often rely on…

Computer Vision and Pattern Recognition · Computer Science 2026-03-17 Haoyu Zhang , Wei Zhai , Yuhang Yang , Yang Cao , Zheng-Jun Zha

Humans possess a remarkable ability to integrate auditory and visual information, enabling a deeper understanding of the surrounding environment. This early fusion of audio and visual cues, demonstrated through cognitive psychology and…

Computer Vision and Pattern Recognition · Computer Science 2023-12-05 Shentong Mo , Pedro Morgado

Various heuristic objectives for modeling hand-object interaction have been proposed in past work. However, due to the lack of a cohesive framework, these objectives often possess a narrow scope of applicability and are limited by their…

Computer Vision and Pattern Recognition · Computer Science 2023-12-27 Shutong Zhang , Yi-Ling Qiao , Guanglei Zhu , Eric Heiden , Dylan Turpin , Jingzhou Liu , Ming Lin , Miles Macklin , Animesh Garg

We revisit the role of texture in monocular 3D hand reconstruction, not as an afterthought for photorealism, but as a dense, spatially grounded cue that can actively support pose and shape estimation. Our observation is simple: even in…

Computer Vision and Pattern Recognition · Computer Science 2025-08-14 Giorgos Karvounas , Nikolaos Kyriazis , Iason Oikonomidis , Georgios Pavlakos , Antonis A. Argyros

Human motion forecasting, with the goal of estimating future human behavior over a period of time, is a fundamental task in many real-world applications. However, existing works typically concentrate on predicting the major joints of the…

Computer Vision and Pattern Recognition · Computer Science 2024-04-05 Pengxiang Ding , Qiongjie Cui , Min Zhang , Mengyuan Liu , Haofan Wang , Donglin Wang

Human-Object Interaction (HOI) modelling captures how humans act upon and relate to objects, typically expressed as <person, action, object> triplets. Existing approaches split into two disjoint families: HOI generation synthesises scenes…

Computer Vision and Pattern Recognition · Computer Science 2026-04-16 Jiun Tian Hoe , Weipeng Hu , Xudong Jiang , Yap-Peng Tan , Chee Seng Chan

Controllable affordance Hand-Object Interaction (HOI) generation has become an increasingly important area of research in computer vision. In HOI generation, the hand grasp generation is a crucial step for effectively controlling the…

Computer Vision and Pattern Recognition · Computer Science 2025-01-28 Ishant , Rongliang Wu , Joo Hwee Lim

We present a novel Hand-pose Embedding Interactive System (HpEIS) as a virtual sensor, which maps users' flexible hand poses to a two-dimensional visual space using a Variational Autoencoder (VAE) trained on a variety of hand poses. HpEIS…

Computer Vision and Pattern Recognition · Computer Science 2024-10-14 Songpei Xu , Xuri Ge , Chaitanya Kaul , Roderick Murray-Smith

3D hand-object pose estimation is the key to the success of many computer vision applications. The main focus of this task is to effectively model the interaction between the hand and an object. To this end, existing works either rely on…

Computer Vision and Pattern Recognition · Computer Science 2023-01-09 Rong Wang , Wei Mao , Hongdong Li