中文
相关论文

相关论文: HOI-PAGE: Zero-Shot Human-Object Interaction Gener…

200 篇论文

Synthesizing realistic human-object interactions (HOI) in video is challenging due to the complex, instance-specific interaction dynamics of both humans and objects. Incorporating controllability in video generation further adds to the…

计算机视觉与模式识别 · 计算机科学 2026-04-09 Wanyue Zhang , Lin Geng Foo , Thabo Beeler , Rishabh Dabral , Christian Theobalt

Hand-object interaction (HOI) reconstruction and synthesis are becoming central to embodied AI and AR/VR. Yet, despite rapid progress, existing HOI generation research remains fragmented across three disjoint tracks: (1) pose-only synthesis…

计算机视觉与模式识别 · 计算机科学 2026-04-03 Mingju Gao , Kaisen Yang , Huan-ang Gao , Bohan Li , Ao Ding , Wenyi Li , Yangcheng Yu , Jinkun Liu , Shaocong Xu , Yike Niu , Haohan Chi , Hao Chen , Hao Tang , Yu Zhang , Li Yi , Hao Zhao

Controllable affordance Hand-Object Interaction (HOI) generation has become an increasingly important area of research in computer vision. In HOI generation, the hand grasp generation is a crucial step for effectively controlling the…

计算机视觉与模式识别 · 计算机科学 2025-01-28 Ishant , Rongliang Wu , Joo Hwee Lim

The Human-Object Interaction (HOI) task explores the dynamic interactions between humans and objects in physical environments, providing essential biomechanical and cognitive-behavioral foundations for fields such as robotics, virtual…

计算机视觉与模式识别 · 计算机科学 2025-08-11 Ruiyan Wang , Lin Zuo , Zonghao Lin , Qiang Wang , Zhengxue Cheng , Rong Xie , Jun Ling , Li Song

We address the task of generating physically accurate and visually faithful 4D Human-Object Interaction (HOI). Given a static 3D human and target object represented as 3D Gaussian Splats (3DGS), our goal is to synthesize dynamic scenes…

计算机视觉与模式识别 · 计算机科学 2026-05-29 Omer Benishu , Gal Fiebelman , Sagie Benaim

Understanding the inherent human knowledge in interacting with a given environment (e.g., affordance) is essential for improving AI to better assist humans. While existing approaches primarily focus on human-object contacts during…

计算机视觉与模式识别 · 计算机科学 2024-07-24 Hyeonwoo Kim , Sookwan Han , Patrick Kwon , Hanbyul Joo

Text-to-video (T2V) generation has made tremendous progress in generating complicated scenes based on texts. However, human-object interaction (HOI) often cannot be precisely generated by current T2V models due to the lack of large-scale…

计算机视觉与模式识别 · 计算机科学 2025-04-01 Kun Liu , Qi Liu , Xinchen Liu , Jie Li , Yongdong Zhang , Jiebo Luo , Xiaodong He , Wu Liu

Compositional Customized Image Generation aims to customize multiple target concepts within generation content, which has gained attention for its wild application. Existing approaches mainly concentrate on the target entity's appearance…

计算机视觉与模式识别 · 计算机科学 2025-08-29 Zhu Xu , Zhaowen Wang , Yuxin Peng , Yang Liu

Most existing Human-Object Interaction~(HOI) Detection methods rely heavily on full annotations with predefined HOI categories, which is limited in diversity and costly to scale further. We aim at advancing zero-shot HOI detection to detect…

计算机视觉与模式识别 · 计算机科学 2022-11-28 Mingrui Wu , Jiaxin Gu , Yunhang Shen , Mingbao Lin , Chao Chen , Xiaoshuai Sun

We present an approach for detecting human-object interactions (HOIs) in images, based on the idea that humans interact with functionally similar objects in a similar manner. The proposed model is simple and efficiently uses the data,…

计算机视觉与模式识别 · 计算机科学 2020-09-03 Ankan Bansal , Sai Saketh Rambhatla , Abhinav Shrivastava , Rama Chellappa

Hand-Object Interaction (HOI) remains a core challenge in digital human video synthesis, where models must generate physically plausible contact and preserve object identity across frames. Although recent HOI reenactment approaches have…

计算机视觉与模式识别 · 计算机科学 2026-03-09 Xuan Huang , Mochu Xiang , Zhelun Shen , Jinbo Wu , Chenming Wu , Chen Zhao , Kaisiyuan Wang , Hang Zhou , Shanshan Liu , Haocheng Feng , Wei He , Jingdong Wang

Recent vision-language-action (VLA) models can generate plausible end-effector motions, yet they often fail in long-horizon, contact-rich tasks because the underlying hand-object interaction (HOI) structure is not explicitly represented. An…

机器人学 · 计算机科学 2026-02-17 Huajian Zeng , Lingyun Chen , Jiaqi Yang , Yuantai Zhang , Fan Shi , Peidong Liu , Xingxing Zuo

While existing image-guided composition methods may help insert a foreground object onto a user-specified region of a background image, achieving natural blending inside the region with the rest of the image unchanged, we observe that these…

计算机视觉与模式识别 · 计算机科学 2025-07-23 Dong Liang , Jinyuan Jia , Yuhao Liu , Rynson W. H. Lau

Fine-grained capturing of 3D HOI boosts human activity understanding and facilitates downstream visual tasks, including action recognition, holistic scene reconstruction, and human motion synthesis. Despite its significance, existing works…

计算机视觉与模式识别 · 计算机科学 2023-12-19 Nan Jiang , Tengyu Liu , Zhexuan Cao , Jieming Cui , Zhiyuan zhang , Yixin Chen , He Wang , Yixin Zhu , Siyuan Huang

We study the problem of precisely swapping objects in videos, with a focus on those interacted with by hands, given one user-provided reference object image. Despite the great advancements that diffusion models have made in video editing…

计算机视觉与模式识别 · 计算机科学 2024-11-12 Zihui Xue , Mi Luo , Changan Chen , Kristen Grauman

Synthesizing realistic human-object interaction (HOI) is essential for 3D computer vision and robotics, underpinning animation and embodied control. Existing approaches often require manually specified intermediate waypoints and place all…

计算机视觉与模式识别 · 计算机科学 2026-04-14 Hwanhee Jung , Seunggwan Lee , Jeongyoon Yoon , SeungHyeon Kim , Giljoo Nam , Qixing Huang , Sangpil Kim

Human-Object Interaction (HOI) video reenactment with realistic motion remains a frontier in expressive digital human creation. Existing approaches primarily handle simple image-plane motion (e.g., in-plane translations), struggling with…

To address key limitations in human-object interaction (HOI) video generation -- specifically the reliance on curated motion data, limited generalization to novel objects/scenarios, and restricted accessibility -- we introduce…

计算机视觉与模式识别 · 计算机科学 2025-06-11 Ziyao Huang , Zixiang Zhou , Juan Cao , Yifeng Ma , Yi Chen , Zejing Rao , Zhiyong Xu , Hongmei Wang , Qin Lin , Yuan Zhou , Qinglin Lu , Fan Tang

This paper introduces the first text-guided work for generating the sequence of hand-object interaction in 3D. The main challenge arises from the lack of labeled data where existing ground-truth datasets are nowhere near generalizable in…

计算机视觉与模式识别 · 计算机科学 2024-04-03 Junuk Cha , Jihyeon Kim , Jae Shin Yoon , Seungryul Baek

Learning-based methods to understand and model hand-object interactions (HOI) require a large amount of high-quality HOI data. One way to create HOI data is to transfer hand poses from a source object to another based on the objects'…

计算机视觉与模式识别 · 计算机科学 2026-01-07 Qiaochu Wang , Chufeng Xiao , Manfred Lau , Hongbo Fu