English
Related papers

Related papers: Imagine2Real: Towards Zero-shot Humanoid-Object In…

200 papers

We introduce D3D-HOI: a dataset of monocular videos with ground truth annotations of 3D object pose, shape and part motion during human-object interactions. Our dataset consists of several common articulated objects captured from diverse…

Computer Vision and Pattern Recognition · Computer Science 2021-08-20 Xiang Xu , Hanbyul Joo , Greg Mori , Manolis Savva

We propose a new dataset and a novel approach to learning hand-object interaction priors for hand and articulated object pose estimation. We first collect a dataset using visual teleoperation, where the human operator can directly play…

Computer Vision and Pattern Recognition · Computer Science 2024-07-30 Zehao Zhu , Jiashun Wang , Yuzhe Qin , Deqing Sun , Varun Jampani , Xiaolong Wang

In this work, we are dedicated to a new task, i.e., hand-object interaction image generation, which aims to conditionally generate the hand-object image under the given hand, object and their interaction status. This task is challenging and…

Computer Vision and Pattern Recognition · Computer Science 2022-11-29 Hezhen Hu , Weilun Wang , Wengang Zhou , Houqiang Li

We tackle the challenging problem of human-object interaction (HOI) detection. Existing methods either recognize the interaction of each human-object pair in isolation or perform joint inference based on complex appearance-based features.…

Computer Vision and Pattern Recognition · Computer Science 2020-08-27 Chen Gao , Jiarui Xu , Yuliang Zou , Jia-Bin Huang

Reconstructing textured 3D human models from a single image is fundamental for AR/VR and digital human applications. However, existing methods mostly focus on single individuals and thus fail in multi-human scenes, where naive composition…

Computer Vision and Pattern Recognition · Computer Science 2026-04-08 Gwanghyun Kim , Junghun James Kim , Suh Yoon Jeon , Jason Park , Se Young Chun

3D human-object interaction (HOI) anticipation aims to predict the future motion of humans and their manipulated objects, conditioned on the historical context. Generally, the articulated humans and rigid objects exhibit different motion…

Computer Vision and Pattern Recognition · Computer Science 2025-08-12 Xiaotong Lin , Tianming Liang , Jian-Fang Hu , Kun-Yu Lin , Yulei Kang , Chunwei Tian , Jianhuang Lai , Wei-Shi Zheng

Human-scene interaction (HSI) generation is crucial for applications in embodied AI, virtual reality, and robotics. Yet, existing methods cannot synthesize interactions in unseen environments such as in-the-wild scenes or reconstructed…

Computer Vision and Pattern Recognition · Computer Science 2025-03-24 Hongjie Li , Hong-Xing Yu , Jiaman Li , Jiajun Wu

While existing image-guided composition methods may help insert a foreground object onto a user-specified region of a background image, achieving natural blending inside the region with the rest of the image unchanged, we observe that these…

Computer Vision and Pattern Recognition · Computer Science 2025-07-23 Dong Liang , Jinyuan Jia , Yuhao Liu , Rynson W. H. Lau

Hand-object interaction(HOI) is the fundamental link between human and environment, yet its dexterous and complex pose significantly challenges for gesture control. Despite significant advances in AI and robotics, enabling machines to…

Robotics · Computer Science 2025-07-11 Yongqi Tian , Xueyu Sun , Haoyuan He , Linji Hao , Ning Ding , Caigui Jiang

Monocular 4D human-object interaction (HOI) reconstruction - recovering a moving human and a manipulated object from a single RGB video - remains challenging due to depth ambiguity and frequent occlusions. Existing methods often rely on…

Computer Vision and Pattern Recognition · Computer Science 2026-03-17 Haoyu Zhang , Wei Zhai , Yuhang Yang , Yang Cao , Zheng-Jun Zha

Enabling humanoid robots to clean rooms has long been a pursued dream within humanoid research communities. However, many tasks require multi-humanoid collaboration, such as carrying large and heavy furniture together. Given the scarcity of…

Robotics · Computer Science 2024-10-31 Jiawei Gao , Ziqin Wang , Zeqi Xiao , Jingbo Wang , Tai Wang , Jinkun Cao , Xiaolin Hu , Si Liu , Jifeng Dai , Jiangmiao Pang

Humans rarely plan whole-body interactions with objects at the level of explicit whole-body movements. High-level intentions, such as affordance, define the goal, while coordinated balance, contact, and manipulation can emerge naturally…

Computer Vision and Pattern Recognition · Computer Science 2026-02-06 Sirui Xu , Samuel Schulter , Morteza Ziyadi , Xialin He , Xiaohan Fei , Yu-Xiong Wang , Liangyan Gui

Hand motion capture data is now relatively easy to obtain, even for complicated grasps; however this data is of limited use without the ability to retarget it onto the hands of a specific character or robot. The target hand may differ…

Graphics · Computer Science 2024-02-08 Arjun S. Lakshmipathy , Jessica K. Hodgins , Nancy S. Pollard

Reconstructing human-object interactions (HOI) from single images is fundamental in computer vision. Existing methods are primarily trained and tested on indoor scenes due to the lack of 3D data, particularly constrained by the object…

Computer Vision and Pattern Recognition · Computer Science 2025-03-21 Boran Wen , Dingbang Huang , Zichen Zhang , Jiahong Zhou , Jianbin Deng , Jingyu Gong , Yulong Chen , Lizhuang Ma , Yong-Lu Li

Multimodal Large Language Models have shown promising capabilities in bridging visual and textual reasoning, yet their reasoning capabilities in Open-Vocabulary Human-Object Interaction (OV-HOI) are limited by cross-modal hallucinations and…

Computer Vision and Pattern Recognition · Computer Science 2026-05-21 Zhenlong Yuan , Yue Wang , Dapeng Zhang , Kejin Cui , Rui Chen , Jing Tang , Lei Sun , Hongwei Yu , Chengxuan Qian , Xiangxiang Chu , Shuo Li , Yuyin Zhou

Understanding humans from LiDAR point clouds is one of the most critical tasks in autonomous driving due to its close relationships with pedestrian safety, yet it remains challenging in the presence of diverse human-object interactions and…

Computer Vision and Pattern Recognition · Computer Science 2026-03-18 Daniel Sungho Jung , Dohee Cho , Kyoung Mu Lee

Presenting high-resolution (HR) human appearance is always critical for the human-centric videos. However, current imagery equipment can hardly capture HR details all the time. Existing super-resolution algorithms barely mitigate the…

Computer Vision and Pattern Recognition · Computer Science 2020-05-28 Guanghan Li , Yaping Zhao , Mengqi Ji , Xiaoyun Yuan , Lu Fang

Training perceptive humanoid locomotion policies that traverse complex terrains with natural gaits remains an open challenge, typically demanding multi-stage training pipelines, adversarial objectives, or extensive real-world calibration.…

Robotics · Computer Science 2026-03-20 Chenxi Han , Shilu He , Yi Cheng , Linqi Ye , Houde Liu

Text-conditioned motion synthesis has made remarkable progress with the emergence of diffusion models. However, the majority of these motion diffusion models are primarily designed for a single character and overlook multi-human…

Computer Vision and Pattern Recognition · Computer Science 2024-11-22 Zhenzhi Wang , Jingbo Wang , Yixuan Li , Dahua Lin , Bo Dai

Interactive visual grounding in Human-Robot Interaction (HRI) is challenging yet practical due to the inevitable ambiguity in natural languages. It requires robots to disambiguate the user input by active information gathering. Previous…

Robotics · Computer Science 2024-02-20 Jie Xu , Hanbo Zhang , Qingyi Si , Yifeng Li , Xuguang Lan , Tao Kong