English
Related papers

Related papers: MaskHOI: Robust 3D Hand-Object Interaction Estimat…

200 papers

Human action recognition is a crucial task for intelligent robotics, particularly within the context of human-robot collaboration research. In self-supervised skeleton-based action recognition, the mask-based reconstruction paradigm learns…

Computer Vision and Pattern Recognition · Computer Science 2025-08-19 Wei Wei , Shaojie Zhang , Yonghao Dang , Jianqin Yin

Human-Object Interaction (HOI) video reenactment with realistic motion remains a frontier in expressive digital human creation. Existing approaches primarily handle simple image-plane motion (e.g., in-plane translations), struggling with…

Computer Vision and Pattern Recognition · Computer Science 2026-03-17 Jinguang Tong , Jinbo Wu , Kaisiyuan Wang , Zhelun Shen , Xuan Huang , Mochu Xiang , Xuesong Li , Yingying Li , Haocheng Feng , Chen Zhao , Hang Zhou , Wei He , Chuong Nguyen , Jingdong Wang , Hongdong Li

In the field of human-object interaction (HOI), detection and generation are two dual tasks that have traditionally been addressed separately, hindering the development of comprehensive interaction understanding. To address this, we propose…

Computer Vision and Pattern Recognition · Computer Science 2025-11-20 Panqi Yang , Haodong Jing , Nanning Zheng , Yongqiang Ma

Human-Object Interaction (HOI) detection is a longstanding computer vision problem concerned with predicting the interaction between humans and objects. Current HOI models rely on a vocabulary of interactions at training and inference time,…

Computer Vision and Pattern Recognition · Computer Science 2026-04-16 Francesco Tonini , Alessandro Conti , Lorenzo Vaquero , Cigdem Beyan , Elisa Ricci

Human hands are highly articulated and versatile at handling objects. Jointly estimating the 3D poses of a hand and the object it manipulates from a monocular camera is challenging due to frequent occlusions. Thus, existing methods often…

Computer Vision and Pattern Recognition · Computer Science 2024-02-28 Haozhe Qi , Chen Zhao , Mathieu Salzmann , Alexander Mathis

In virtual Hand-Object Interaction (HOI) scenarios, the authenticity of the hand's deformation is important to immersive experience, such as natural manipulation or tactile feedback. Unrealistic deformation arises from simplified hand…

In general, hand pose estimation aims to improve the robustness of model performance in the real-world scenes. However, it is difficult to enhance the robustness since existing datasets are obtained in restricted environments to annotate 3D…

Computer Vision and Pattern Recognition · Computer Science 2023-10-20 Bosang Kim , Jonghyun Kim , Hyotae Lee , Lanying Jin , Jeongwon Ha , Dowoo Kwon , Jungpyo Kim , Wonhyeok Im , KyungMin Jin , Jungho Lee

Executing reliable Humanoid-Object Interaction (HOI) tasks for humanoid robots is hindered by the lack of generalized control interfaces and robust closed-loop perception mechanisms. In this work, we introduce Perceptive Root-guided…

Robotics · Computer Science 2026-03-03 Yuhang Lin , Jiyuan Shi , Dewei Wang , Jipeng Kong , Yong Liu , Chenjia Bai , Xuelong Li

Recent vision-language-action (VLA) models can generate plausible end-effector motions, yet they often fail in long-horizon, contact-rich tasks because the underlying hand-object interaction (HOI) structure is not explicitly represented. An…

Robotics · Computer Science 2026-02-17 Huajian Zeng , Lingyun Chen , Jiaqi Yang , Yuantai Zhang , Fan Shi , Peidong Liu , Xingxing Zuo

With the diversification of human-object interaction (HOI) applications and the success of capturing human meshes, HOI reconstruction has gained widespread attention. Existing mainstream HOI reconstruction methods often rely on explicitly…

Computer Vision and Pattern Recognition · Computer Science 2025-03-11 Zhenrong Wang , Qi Zheng , Sihan Ma , Maosheng Ye , Yibing Zhan , Dongjiang Li

Self-supervised learning guided by masked image modelling, such as Masked AutoEncoder (MAE), has attracted wide attention for pretraining vision transformers in remote sensing. However, MAE tends to excessively focus on pixel details,…

Computer Vision and Pattern Recognition · Computer Science 2023-10-31 Yi Wang , Hugo Hernández Hernández , Conrad M Albrecht , Xiao Xiang Zhu

Masked Autoencoders (MAE) have shown great potentials in self-supervised pre-training for language and 2D image transformers. However, it still remains an open question on how to exploit masked autoencoding for learning 3D representations…

Computer Vision and Pattern Recognition · Computer Science 2022-10-17 Renrui Zhang , Ziyu Guo , Rongyao Fang , Bin Zhao , Dong Wang , Yu Qiao , Hongsheng Li , Peng Gao

Current video-based Masked Autoencoders (MAEs) primarily focus on learning effective spatiotemporal representations from a visual perspective, which may lead the model to prioritize general spatial-temporal patterns but often overlook…

Computer Vision and Pattern Recognition · Computer Science 2025-02-13 Shihab Aaqil Ahamed , Malitha Gunawardhana , Liel David , Michael Sidorov , Daniel Harari , Muhammad Haris Khan

This paper addresses the task of 3D pose estimation for a hand interacting with an object from a single image observation. When modeling hand-object interaction, previous works mainly exploit proximity cues, while overlooking the dynamical…

Computer Vision and Pattern Recognition · Computer Science 2023-12-12 Rong Wang , Wei Mao , Hongdong Li

The lifting-based methods have dominated monocular 3D human pose estimation by leveraging detected 2D poses as intermediate representations. The 2D component of the final 3D human pose benefits from the detected 2D poses, whereas its depth…

Computer Vision and Pattern Recognition · Computer Science 2025-12-19 Mengyuan Liu , Jiajie Liu , Jinyan Zhang , Wenhao Li , Junsong Yuan

Generalized robots must learn from diverse, large-scale human-object interactions (HOI) to operate robustly in the real world. Monocular internet videos offer a nearly limitless and readily available source of data, capturing an…

Computer Vision and Pattern Recognition · Computer Science 2026-03-20 Boran Wen , Ye Lu , Sirui Wang , Keyan Wan , Jiahong Zhou , Junxuan Liang , Xinpeng Liu , Bang Xiao , Ruiyang Liu , Yong-Lu Li

We propose a multimodal, physically grounded approach for metric-scale amodal object reconstruction and pose estimation under severe hand occlusion. Unlike prior occlusion-aware 3D generation methods that rely only on vision, we leverage…

Computer Vision and Pattern Recognition · Computer Science 2026-04-13 Gabriele Mario Caddeo , Pasquale Marra , Lorenzo Natale

We present HOIGaze - a novel learning-based approach for gaze estimation during hand-object interactions (HOI) in extended reality (XR). HOIGaze addresses the challenging HOI setting by building on one key insight: The eye, hand, and head…

Computer Vision and Pattern Recognition · Computer Science 2025-04-29 Zhiming Hu , Daniel Haeufle , Syn Schmitt , Andreas Bulling

We introduce Robowheel, a data engine that converts human hand object interaction (HOI) videos into training-ready supervision for cross morphology robotic learning. From monocular RGB or RGB-D inputs, we perform high precision HOI…

Egocentric hand-object motion generation is crucial for immersive AR/VR and robotic imitation but remains challenging due to unstable viewpoints, self-occlusions, perspective distortion, and noisy ego-motion. Existing methods rely on…

Computer Vision and Pattern Recognition · Computer Science 2025-05-23 Bohan Zhou , Yi Zhan , Zhongbin Zhang , Zongqing Lu