English
Related papers

Related papers: Affordance-Guided Diffusion Prior for 3D Hand Reco…

200 papers

Diffusion models when conditioned on text prompts, generate realistic-looking images with intricate details. But most of these pre-trained models fail to generate accurate images when it comes to human features like hands, teeth, etc. We…

Computer Vision and Pattern Recognition · Computer Science 2023-12-27 Gurusha Juneja , Sukrit Kumar

Diffusion models have achieved remarkable success in generating realistic images but suffer from generating accurate human hands, such as incorrect finger counts or irregular shapes. This difficulty arises from the complex task of learning…

Computer Vision and Pattern Recognition · Computer Science 2024-08-19 Wenquan Lu , Yufei Xu , Jing Zhang , Chaoyue Wang , Dacheng Tao

Pose and body shape editing in a human image has received increasing attention. However, current methods often struggle with dataset biases and deteriorate realism and the person's identity when users make large edits. We propose a one-shot…

Computer Vision and Pattern Recognition · Computer Science 2024-01-09 Yuta Okuyama , Yuki Endo , Yoshihiro Kanamori

We propose a multimodal, physically grounded approach for metric-scale amodal object reconstruction and pose estimation under severe hand occlusion. Unlike prior occlusion-aware 3D generation methods that rely only on vision, we leverage…

Computer Vision and Pattern Recognition · Computer Science 2026-04-13 Gabriele Mario Caddeo , Pasquale Marra , Lorenzo Natale

Diffusion-based policies have shown impressive performance in robotic manipulation tasks while struggling with out-of-domain distributions. Recent efforts attempted to enhance generalization by improving the visual feature encoding for…

Robotics · Computer Science 2025-03-21 Shijie Wu , Yihang Zhu , Yunao Huang , Kaizhen Zhu , Jiayuan Gu , Jingyi Yu , Ye Shi , Jingya Wang

Existing multi-view 3D object reconstruction methods heavily rely on sufficient overlap between input views, where occlusions and sparse coverage in practice frequently yield severe reconstruction incompleteness. Recent advancements in…

Computer Vision and Pattern Recognition · Computer Science 2025-10-28 Jiahao Chang , Chongjie Ye , Yushuang Wu , Yuantao Chen , Yidan Zhang , Zhongjin Luo , Chenghong Li , Yihao Zhi , Xiaoguang Han

Predicting and generating human hand grasp over objects is critical for animation and robotic tasks. In this work, we focus on generating both the hand and objects in a grasp by a single diffusion model. Our proposed Joint Hand-Object…

Computer Vision and Pattern Recognition · Computer Science 2026-01-28 Jinkun Cao , Jingyuan Liu , Kris Kitani , Yi Zhou

Textured 3D morphing creates smooth and plausible interpolation sequences between two 3D objects, focusing on transitions in both shape and texture. This is important for creative applications like visual effects in filmmaking. Previous…

Computer Vision and Pattern Recognition · Computer Science 2025-02-21 Songlin Yang , Yushi Lan , Honghua Chen , Xingang Pan

A core problem of Embodied AI is to learn object manipulation from observation, as humans do. To achieve this, it is important to localize 3D object affordance areas through observation such as images (3D affordance grounding) and…

Computer Vision and Pattern Recognition · Computer Science 2026-04-02 Xinhang Wan , Dongqiang Gou , Xinwang Liu , En Zhu , Xuming He

Recent generative models can synthesize high-quality images, but they often fail to generate humans interacting with objects using their hands. This arises mostly from the model's misunderstanding of such interactions and the hardships of…

Computer Vision and Pattern Recognition · Computer Science 2025-12-01 Patrick Kwon , Chen Chen , Hanbyul Joo

Reconstructing two-hand interactions from a single image is a challenging problem due to ambiguities that stem from projective geometry and heavy occlusions. Existing methods are designed to estimate only a single pose, despite the fact…

Computer Vision and Pattern Recognition · Computer Science 2022-10-05 Jiayi Wang , Diogo Luvizon , Franziska Mueller , Florian Bernard , Adam Kortylewski , Dan Casas , Christian Theobalt

Lexical semantics and cognitive science point to affordances (i.e. the actions that objects support) as critical for understanding and representing nouns and verbs. However, study of these semantic features has not yet been integrated with…

Computation and Language · Computer Science 2022-07-07 Jack Merullo , Dylan Ebert , Carsten Eickhoff , Ellie Pavlick

Current digital human studies focusing on lip-syncing and body movement are no longer sufficient to meet the growing industrial demand, while human video generation techniques that support interacting with real-world environments (e.g.,…

Computer Vision and Pattern Recognition · Computer Science 2025-03-26 Yingying Fan , Quanwei Yang , Kaisiyuan Wang , Hang Zhou , Yingying Li , Haocheng Feng , Errui Ding , Yu Wu , Jingdong Wang

We propose G-HOP, a denoising diffusion based generative prior for hand-object interactions that allows modeling both the 3D object and a human hand, conditioned on the object category. To learn a 3D spatial diffusion model that can capture…

Computer Vision and Pattern Recognition · Computer Science 2024-04-19 Yufei Ye , Abhinav Gupta , Kris Kitani , Shubham Tulsiani

Understanding how humans would behave during hand-object interaction is vital for applications in service robot manipulation and extended reality. To achieve this, some recent works have been proposed to simultaneously forecast hand…

Computer Vision and Pattern Recognition · Computer Science 2025-11-17 Junyi Ma , Jingyi Xu , Xieyuanli Chen , Hesheng Wang

Affordance learning considers the interaction opportunities for an actor in the scene and thus has wide application in scene understanding and intelligent robotics. In this paper, we focus on contextual affordance learning, i.e., using…

Computer Vision and Pattern Recognition · Computer Science 2023-08-07 Jieteng Yao , Junjie Chen , Li Niu , Bin Sheng

Object affordance is an important concept in hand-object interaction, providing information on action possibilities based on human motor capacity and objects' physical property thus benefiting tasks such as action anticipation and robot…

Computer Vision and Pattern Recognition · Computer Science 2023-02-13 Zecheng Yu , Yifei Huang , Ryosuke Furuta , Takuma Yagi , Yusuke Goutsu , Yoichi Sato

Diffusion-based methods have achieved significant successes in T2I generation, providing realistic images from text prompts. Despite their capabilities, these models face persistent challenges in generating realistic human hands, often…

Computer Vision and Pattern Recognition · Computer Science 2025-03-12 Taehyeon Eum , Jieun Choi , Tae-Kyun Kim

Full 3D estimation of human pose from a single image remains a challenging task despite many recent advances. In this paper, we explore the hypothesis that strong prior information about scene geometry can be used to improve pose estimation…

Computer Vision and Pattern Recognition · Computer Science 2021-12-10 Zhe Wang , Liyan Chen , Shaurya Rathore , Daeyun Shin , Charless Fowlkes

Recently, 3D hand reconstruction has gained more attention in human-computer cooperation, especially for hand-object interaction scenario. However, it still remains huge challenge due to severe hand-occlusion caused by interaction, which…

Computer Vision and Pattern Recognition · Computer Science 2024-03-05 Feng Shuang , Wenbo He , Shaodong Li