English
Related papers

Related papers: StructBiHOI: Structured Articulation Modeling for …

200 papers

Achieving realistic simulations of humans interacting with a wide range of objects has long been a fundamental goal. Extending physics-based motion imitation to complex human-object interactions (HOIs) is challenging due to intricate…

Computer Vision and Pattern Recognition · Computer Science 2026-02-03 Sirui Xu , Hung Yu Ling , Yu-Xiong Wang , Liang-Yan Gui

Text-conditioned human motion generation has experienced significant advancements with diffusion models trained on extensive motion capture data and corresponding textual annotations. However, extending such success to 3D dynamic…

Computer Vision and Pattern Recognition · Computer Science 2026-02-03 Sirui Xu , Ziyin Wang , Yu-Xiong Wang , Liang-Yan Gui

Autonomous robotic assembly of interlocking bricks demands seamless integration of long-horizon task reasoning, spatial grounding, and fine-grained manipulation. This paper presents BrickCraft, a compositional framework designed for…

Robotics · Computer Science 2026-05-11 Jichuan Yu , Bowei Li , Zhenran Tang , Guanxing Lu , Chuxiong Hu , Ruixuan Liu , Changliu Liu

Dexterous hands enable concurrent prehensile and nonprehensile manipulation, such as holding one object while interacting with another, a capability essential for everyday tasks yet underexplored in robotics. Learning such long-horizon,…

Robotics · Computer Science 2026-03-17 Hao Jiang , Yue Wu , Yue Wang , Gaurav S. Sukhatme , Daniel Seita

Human-Object Interaction (HOI) detection is a challenging computer vision task that requires visual models to address the complex interactive relationship between humans and objects and predict HOI triplets. Despite the challenges posed by…

Computer Vision and Pattern Recognition · Computer Science 2023-09-19 Yichao Cao , Qingfei Tang , Feng Yang , Xiu Su , Shan You , Xiaobo Lu , Chang Xu

Human-object interaction (HOI) video generation has garnered increasing attention due to its promising applications in digital humans, e-commerce, advertising, and robotics imitation learning. However, existing methods face two critical…

Computer Vision and Pattern Recognition · Computer Science 2026-03-27 Bangya Liu , Xinyu Gong , Zelin Zhao , Ziyang Song , Yulei Lu , Suhui Wu , Jun Zhang , Suman Banerjee , Hao Zhang

Recent studies on 3D hand reconstruction have demonstrated the effectiveness of synthetic training data to improve estimation performance. However, most methods rely on game engines to synthesize hand images, which often lack diversity in…

Computer Vision and Pattern Recognition · Computer Science 2026-03-03 Zhuoran Zhao , Xianghao Kong , Linlin Yang , Zheng Wei , Pan Hui , Anyi Rao

Humans are highly adaptable, swiftly switching between different modes to progressively handle different tasks, situations and contexts. In Human-object interaction (HOI) activities, these modes can be attributed to two mechanisms: (1) the…

Computer Vision and Pattern Recognition · Computer Science 2023-07-25 Hung Tran , Vuong Le , Svetha Venkatesh , Truyen Tran

Object grasping is an important ability required for various robot tasks. In particular, tasks that require precise force adjustments during operation, such as grasping an unknown object or using a grasped tool, are difficult for humans to…

Robotics · Computer Science 2024-01-22 Koki Yamane , Sho Sakaino , Toshiaki Tsuji

This paper presents a novel approach to generating the 3D motion of a human interacting with a target object, with a focus on solving the challenge of synthesizing long-range and diverse motions, which could not be fulfilled by existing…

Computer Vision and Pattern Recognition · Computer Science 2023-10-04 Huaijin Pi , Sida Peng , Minghui Yang , Xiaowei Zhou , Hujun Bao

Two-hand reconstruction from monocular images is hampered by complex poses and severe occlusions, which often cause interaction misalignment and two-hand penetration. We address this by decoupling the problem into 2D structural alignment…

Computer Vision and Pattern Recognition · Computer Science 2026-03-10 Gaoge Han , Yongkang Cheng , Zhe Chen , Shaoli Huang , Tongliang Liu

Recent world-model-based Vision-Language-Action (VLA) architectures have improved robotic manipulation through predictive visual foresight. However, dense future prediction introduces visual redundancy and accumulates errors, causing…

Robotics · Computer Science 2026-03-16 Minghao Jin , Mozheng Liao , Mingfei Han , Zhihui Li , Xiaojun Chang

Modeling human-object interactions (HOI) from an egocentric perspective is a critical yet challenging task, particularly when relying on sparse signals from wearable devices like smart glasses and watches. We present ECHO, the first unified…

Computer Vision and Pattern Recognition · Computer Science 2026-03-17 Ilya A. Petrov , Vladimir Guzov , Riccardo Marin , Emre Aksan , Xu Chen , Daniel Cremers , Thabo Beeler , Gerard Pons-Moll

Generating realistic 3D Human-Human Interaction (HHI) requires coherent modeling of the physical plausibility of the agents and their interaction semantics. Existing methods compress all motion information into a single latent…

Computer Vision and Pattern Recognition · Computer Science 2026-03-03 Zichen Geng , Zeeshan Hayder , Bo Miao , Jian Liu , Wei Liu , Ajmal Mian

Human-Object Interaction Detection (HOI-DET) aims to localize human-object pairs and identify their interactive relationships. To aggregate contextual cues, existing methods typically propagate information across all detected entities via…

Computer Vision and Pattern Recognition · Computer Science 2025-10-22 Jiajun Hong , Jianan Wei , Wenguan Wang

We present a generative model to synthesize 3D shapes as sets of handles -- lightweight proxies that approximate the original 3D shape -- for applications in interactive editing, shape parsing, and building compact 3D representations. Our…

Computer Vision and Pattern Recognition · Computer Science 2020-04-08 Matheus Gadelha , Giorgio Gori , Duygu Ceylan , Radomir Mech , Nathan Carr , Tamy Boubekeur , Rui Wang , Subhransu Maji

Long-horizon manipulation remains challenging for vision-language-action (VLA) policies: real tasks are multi-step, progress-dependent, and brittle to compounding execution errors. We present LoHo-Manip, a modular framework that scales…

Robotics · Computer Science 2026-04-24 Isabella Liu , An-Chieh Cheng , Rui Yan , Geng Chen , Ri-Zhao Qiu , Xueyan Zou , Sha Yi , Hongxu Yin , Xiaolong Wang , Sifei Liu

Recently, 3D hand reconstruction has gained more attention in human-computer cooperation, especially for hand-object interaction scenario. However, it still remains huge challenge due to severe hand-occlusion caused by interaction, which…

Computer Vision and Pattern Recognition · Computer Science 2024-03-05 Feng Shuang , Wenbo He , Shaodong Li

To address key limitations in human-object interaction (HOI) video generation -- specifically the reliance on curated motion data, limited generalization to novel objects/scenarios, and restricted accessibility -- we introduce…

Computer Vision and Pattern Recognition · Computer Science 2025-06-11 Ziyao Huang , Zixiang Zhou , Juan Cao , Yifeng Ma , Yi Chen , Zejing Rao , Zhiyong Xu , Hongmei Wang , Qin Lin , Yuan Zhou , Qinglin Lu , Fan Tang

Human-Object Interaction (HOI) detection is a fundamental task in high-level human-centric scene understanding. We propose PhraseHOI, containing a HOI branch and a novel phrase branch, to leverage language prior and improve relation…

Computer Vision and Pattern Recognition · Computer Science 2022-01-19 Zhimin Li , Cheng Zou , Yu Zhao , Boxun Li , Sheng Zhong