English
Related papers

Related papers: THOM: Generating Physically Plausible Hand-Object …

200 papers

3D hand-object interaction data is scarce due to the hardware constraints in scaling up the data collection process. In this paper, we propose HOIDiffusion for generating realistic and diverse 3D hand-object interaction data. Our model is a…

Computer Vision and Pattern Recognition · Computer Science 2024-03-19 Mengqi Zhang , Yang Fu , Zheng Ding , Sifei Liu , Zhuowen Tu , Xiaolong Wang

Generating realistic 3D hand-object interactions (HOI) is a fundamental challenge in computer vision and robotics, requiring both temporal coherence and high-fidelity physical plausibility. Existing methods remain limited in their ability…

Computer Vision and Pattern Recognition · Computer Science 2026-04-14 Zerui Chen , Rolandos Alexandros Potamias , Shizhe Chen , Jiankang Deng , Cordelia Schmid , Stefanos Zafeiriou

We address the task of generating physically accurate and visually faithful 4D Human-Object Interaction (HOI). Given a static 3D human and target object represented as 3D Gaussian Splats (3DGS), our goal is to synthesize dynamic scenes…

Computer Vision and Pattern Recognition · Computer Science 2026-05-29 Omer Benishu , Gal Fiebelman , Sagie Benaim

Generating physically realistic humanoid-object interactions (HOI) is a fundamental challenge in robotics. Existing HOI generation approaches, such as diffusion-based models, often suffer from artifacts such as implausible contacts,…

Robotics · Computer Science 2025-08-21 Yuhang Lin , Yijia Xie , Jiahong Xie , Yuehao Huang , Ruoyu Wang , Jiajun Lv , Yukai Ma , Xingxing Zuo

We present HOIDiNi, a text-driven diffusion framework for synthesizing realistic and plausible human-object interaction (HOI). HOI generation is extremely challenging since it induces strict contact accuracies alongside a diverse motion…

Computer Vision and Pattern Recognition · Computer Science 2025-10-22 Roey Ron , Guy Tevet , Haim Sawdayee , Amit H. Bermano

We introduce Amortized Text-to-Mesh (AToM), a feed-forward text-to-mesh framework optimized across multiple text prompts simultaneously. In contrast to existing text-to-3D methods that often entail time-consuming per-prompt optimization and…

For designing a wide range of everyday objects, the design process should be aware of both the human body and the underlying semantics of the design specification. However, these two objectives present significant challenges to the current…

Computer Vision and Pattern Recognition · Computer Science 2025-03-31 Michelle Guo , Mia Tang , Hannah Cha , Ruohan Zhang , C. Karen Liu , Jiajun Wu

We address the challenging task of text-driven 3D human-object interaction (HOI) motion generation. Existing methods primarily rely on a direct text-to-HOI mapping, which suffers from three key limitations due to the significant…

Computer Vision and Pattern Recognition · Computer Science 2026-02-12 Yin Wang , Ziyao Zhang , Zhiying Leng , Haitian Liu , Frederick W. B. Li , Mu Li , Xiaohui Liang

Reconstructing human-object interactions (HOI) from single images is fundamental in computer vision. Existing methods are primarily trained and tested on indoor scenes due to the lack of 3D data, particularly constrained by the object…

Computer Vision and Pattern Recognition · Computer Science 2025-03-21 Boran Wen , Dingbang Huang , Zichen Zhang , Jiahong Zhou , Jianbin Deng , Jingyu Gong , Yulong Chen , Lizhuang Ma , Yong-Lu Li

Advances in 3D generative AI have enabled the creation of physical objects from text prompts, but challenges remain in creating objects involving multiple component types. We present a pipeline that integrates 3D generative AI with…

Modeling and generating 3D human-object interactions from text is crucial for applications in AR, XR, and gaming. Existing approaches often rely on score distillation from text-to-image models, but their results suffer from the Janus…

Computer Vision and Pattern Recognition · Computer Science 2026-03-13 Agniv Sharma , Xianghui Xie , Tom Fischer , Eddy Ilg , Gerard Pons-Moll

Text-conditioned human motion generation has experienced significant advancements with diffusion models trained on extensive motion capture data and corresponding textual annotations. However, extending such success to 3D dynamic…

Computer Vision and Pattern Recognition · Computer Science 2026-02-03 Sirui Xu , Ziyin Wang , Yu-Xiong Wang , Liang-Yan Gui

In this work, we are dedicated to a new task, i.e., hand-object interaction image generation, which aims to conditionally generate the hand-object image under the given hand, object and their interaction status. This task is challenging and…

Computer Vision and Pattern Recognition · Computer Science 2022-11-29 Hezhen Hu , Weilun Wang , Wengang Zhou , Houqiang Li

We present HOI-PAGE, a new approach that prioritizes part-level affordance reasoning to generate high-fidelity 4D human-object interactions (HOIs) from text prompts in a zero-shot fashion. In contrast to prior works that focus on global,…

Graphics · Computer Science 2026-05-20 Lei Li , Angela Dai

Controllable affordance Hand-Object Interaction (HOI) generation has become an increasingly important area of research in computer vision. In HOI generation, the hand grasp generation is a crucial step for effectively controlling the…

Computer Vision and Pattern Recognition · Computer Science 2025-01-28 Ishant , Rongliang Wu , Joo Hwee Lim

We propose a novel task of text-controlled human object interaction generation in 3D scenes with movable objects. Existing human-scene interaction datasets suffer from insufficient interaction categories and typically only consider…

Computer Vision and Pattern Recognition · Computer Science 2025-09-30 Xinhao Cai , Minghang Zheng , Xin Jin , Yang Liu

Human-object interaction (HOI) synthesis is crucial for creating immersive and realistic experiences for applications such as virtual reality. Existing methods often rely on simplified object representations, such as the object's centroid…

Computer Vision and Pattern Recognition · Computer Science 2025-03-27 Mengqing Xue , Yifei Liu , Ling Guo , Shaoli Huang , Changxing Ding

Digital human motion synthesis is a vibrant research field with applications in movies, AR/VR, and video games. Whereas methods were proposed to generate natural and realistic human motions, most only focus on modeling humans and largely…

Computer Vision and Pattern Recognition · Computer Science 2023-11-07 Quanzhou Li , Jingbo Wang , Chen Change Loy , Bo Dai

Generating high-fidelity full-body human interactions with dynamic objects and static scenes remains a critical challenge in computer graphics and animation. Existing methods for human-object interaction often neglect scene context, leading…

Computer Vision and Pattern Recognition · Computer Science 2025-11-25 Wei Yao , Yunlian Sun , Hongwen Zhang , Yebin Liu , Jinhui Tang

Human-object interaction (HOI) synthesis is crucial for applications in animation, simulation, and robotics. However, existing approaches either rely on expensive motion capture data or require manual reward engineering, limiting their…

Computer Vision and Pattern Recognition · Computer Science 2026-03-05 Zekai Deng , Ye Shi , Kaiyang Ji , Lan Xu , Shaoli Huang , Jingya Wang