English
Related papers

Related papers: HOIDiffusion: Generating Realistic 3D Hand-Object …

200 papers

Diffusion models currently achieve state-of-the-art performance for both conditional and unconditional image generation. However, so far, image diffusion models do not support tasks required for 3D understanding, such as view-consistent 3D…

Computer Vision and Pattern Recognition · Computer Science 2024-02-22 Titas Anciukevičius , Zexiang Xu , Matthew Fisher , Paul Henderson , Hakan Bilen , Niloy J. Mitra , Paul Guerrero

We present a method for teaching machines to understand and model the underlying spatial common sense of diverse human-object interactions in 3D in a self-supervised way. This is a challenging task, as there exist specific manifolds of the…

Computer Vision and Pattern Recognition · Computer Science 2023-09-06 Sookwan Han , Hanbyul Joo

Hand-object interaction(HOI) is the fundamental link between human and environment, yet its dexterous and complex pose significantly challenges for gesture control. Despite significant advances in AI and robotics, enabling machines to…

Robotics · Computer Science 2025-07-11 Yongqi Tian , Xueyu Sun , Haoyuan He , Linji Hao , Ning Ding , Caigui Jiang

In e-commerce and digital marketing, generating high-fidelity human-product demonstration videos is important for effective product presentation. However, most existing frameworks either fail to preserve the identities of both humans and…

Computer Vision and Pattern Recognition · Computer Science 2025-08-28 Lizhen Wang , Zhurong Xia , Tianshu Hu , Pengrui Wang , Pengfei Wei , Zerong Zheng , Ming Zhou , Yuan Zhang , Mingyuan Gao

We present 3DiFACE, a novel method for personalized speech-driven 3D facial animation and editing. While existing methods deterministically predict facial animations from speech, they overlook the inherent one-to-many relationship between…

Computer Vision and Pattern Recognition · Computer Science 2023-12-05 Balamurugan Thambiraja , Sadegh Aliakbarian , Darren Cosker , Justus Thies

Generating human grasping poses that accurately reflect both object geometry and user-specified interaction semantics is essential for natural hand-object interactions in AR/VR and embodied AI. However, existing semantic grasping approaches…

Robotics · Computer Science 2026-03-31 Xiaofei Wu , Yi Zhang , Yumeng Liu , Yuexin Ma , Yujiao Shi , Xuming He

Generating realistic human-object interaction (HOI) animations remains challenging because it requires jointly modeling dynamic human actions and diverse object geometries. Prior diffusion-based approaches often rely on hand-crafted contact…

Computer Vision and Pattern Recognition · Computer Science 2026-03-27 Ziyin Wang , Sirui Xu , Chuan Guo , Bing Zhou , Jiangshan Gong , Jian Wang , Yu-Xiong Wang , Liang-Yan Gui

Hands are central to interacting with our surroundings and conveying gestures, making their inclusion essential for full-body motion synthesis. Despite this, existing human motion synthesis methods fall short: some ignore hand motions…

Computer Vision and Pattern Recognition · Computer Science 2026-01-08 Enes Duran , Nikos Athanasiou , Muhammed Kocabas , Michael J. Black , Omid Taheri

Generating multi-view human images from a single view is a complex and significant challenge. Although recent advancements in multi-view object generation have shown impressive results with diffusion models, novel view synthesis for humans…

Computer Vision and Pattern Recognition · Computer Science 2024-12-05 Yu Feng , Shunsi Zhang , Jian Shu , Hanfeng Zhao , Guoliang Pang , Chi Zhang , Hao Wang

This paper presents a novel approach to generating the 3D motion of a human interacting with a target object, with a focus on solving the challenge of synthesizing long-range and diverse motions, which could not be fulfilled by existing…

Computer Vision and Pattern Recognition · Computer Science 2023-10-04 Huaijin Pi , Sida Peng , Minghui Yang , Xiaowei Zhou , Hujun Bao

We address the challenging task of text-driven 3D human-object interaction (HOI) motion generation. Existing methods primarily rely on a direct text-to-HOI mapping, which suffers from three key limitations due to the significant…

Computer Vision and Pattern Recognition · Computer Science 2026-02-12 Yin Wang , Ziyao Zhang , Zhiying Leng , Haitian Liu , Frederick W. B. Li , Mu Li , Xiaohui Liang

Human-object-scene interactions (HOSI) generation has broad applications in embodied AI, simulation, and animation. Unlike human-object interaction (HOI) and human-scene interaction (HSI), HOSI generation requires reasoning over dynamic…

Computer Vision and Pattern Recognition · Computer Science 2026-04-07 Yude Zou , Junji Gong , Xing Gao , Zixuan Li , Tianxing Chen , Guanjie Zheng

In this work, we explore the role of synthetic data in improving the detection of Hand-Object Interactions from egocentric images. Through extensive experimentation and comparative analysis on VISOR, EgoHOS, and ENIGMA-51 datasets, our…

Computer Vision and Pattern Recognition · Computer Science 2026-04-01 Rosario Leonardi , Antonino Furnari , Francesco Ragusa , Giovanni Maria Farinella

Recent progress in diffusion models has significantly advanced the field of human image animation. While existing methods can generate temporally consistent results for short or regular motions, significant challenges remain, particularly…

Computer Vision and Pattern Recognition · Computer Science 2025-12-29 Shen Zheng , Jiaran Cai , Yuansheng Guan , Shenneng Huang , Xingpei Ma , Junjie Cao , Hanfeng Zhao , Qiang Zhang , Shunsi Zhang , Xiao-Ping Zhang

Diffusion models have made breakthroughs in 3D generation tasks. Current 3D diffusion models focus on reconstructing target shape from images or a set of partial observations. While excelling in global context understanding, they struggle…

Computer Vision and Pattern Recognition · Computer Science 2025-05-20 Yuanbo Wang , Zhaoxuan Zhang , Jiajin Qiu , Dilong Sun , Zhengyu Meng , Xiaopeng Wei , Xin Yang

The success of image generative models has enabled us to build methods that can edit images based on text or other user input. However, these methods are bespoke, imprecise, require additional information, or are limited to only 2D image…

Computer Vision and Pattern Recognition · Computer Science 2025-01-03 Rahul Sajnani , Jeroen Vanbaar , Jie Min , Kapil Katyal , Srinath Sridhar

Reconstructing hand-held objects in 3D from monocular images remains a significant challenge in computer vision. Most existing approaches rely on implicit 3D representations, which produce overly smooth reconstructions and are…

Computer Vision and Pattern Recognition · Computer Science 2025-07-22 Zerui Chen , Rolandos Alexandros Potamias , Shizhe Chen , Cordelia Schmid

We introduce FabricDiffusion, a method for transferring fabric textures from a single clothing image to 3D garments of arbitrary shapes. Existing approaches typically synthesize textures on the garment surface through 2D-to-3D texture…

Computer Vision and Pattern Recognition · Computer Science 2024-10-03 Cheng Zhang , Yuanhao Wang , Francisco Vicente Carrasco , Chenglei Wu , Jinlong Yang , Thabo Beeler , Fernando De la Torre

Reconstructing hand-held objects from monocular RGB images is an appealing yet challenging task. In this task, contacts between hands and objects provide important cues for recovering the 3D geometry of the hand-held objects. Though recent…

Computer Vision and Pattern Recognition · Computer Science 2024-01-17 Junxing Hu , Hongwen Zhang , Zerui Chen , Mengcheng Li , Yunlong Wang , Yebin Liu , Zhenan Sun

Predicting human motion plays a crucial role in ensuring a safe and effective human-robot close collaboration in intelligent remanufacturing systems of the future. Existing works can be categorized into two groups: those focusing on…

Robotics · Computer Science 2023-08-01 Sibo Tian , Minghui Zheng , Xiao Liang