中文
相关论文

相关论文: MOHO: Learning Single-view Hand-held Object Recons…

200 篇论文

In this paper, we introduce ManiVideo, a novel method for generating consistent and temporally coherent bimanual hand-object manipulation videos from given motion sequences of hands and objects. The core idea of ManiVideo is the…

计算机视觉与模式识别 · 计算机科学 2024-12-24 Youxin Pang , Ruizhi Shao , Jiajun Zhang , Hanzhang Tu , Yun Liu , Boyao Zhou , Hongwen Zhang , Yebin Liu

Estimating the 3D pose of hand and potential hand-held object from monocular images is a longstanding challenge. Yet, existing methods are specialized, focusing on either bare-hand or hand interacting with object. No method can flexibly…

计算机视觉与模式识别 · 计算机科学 2025-03-18 Yinqiao Wang , Hao Xu , Pheng-Ann Heng , Chi-Wing Fu

Our work aims to reconstruct hand-object interactions from a single-view image, which is a fundamental but ill-posed task. Unlike methods that reconstruct from videos, multi-view images, or predefined 3D templates, single-view…

计算机视觉与模式识别 · 计算机科学 2025-12-22 Yumeng Liu , Xiaoxiao Long , Zemin Yang , Yuan Liu , Marc Habermann , Christian Theobalt , Yuexin Ma , Wenping Wang

Reconstructing hand-held objects from monocular RGB images is an appealing yet challenging task. In this task, contacts between hands and objects provide important cues for recovering the 3D geometry of the hand-held objects. Though recent…

计算机视觉与模式识别 · 计算机科学 2024-01-17 Junxing Hu , Hongwen Zhang , Zerui Chen , Mengcheng Li , Yunlong Wang , Yebin Liu , Zhenan Sun

Monocular 3D human reconstruction in real-world scenarios remains highly challenging due to frequent occlusions from surrounding objects, people, or image truncation. Such occlusions lead to missing geometry and unreliable appearance cues,…

计算机视觉与模式识别 · 计算机科学 2026-03-17 Yuanwang Yang , Hongliang Liu , Muxin Zhang , Nan Ma , Jingyu Yang , Yu-Kun Lai , Kun Li

Self-supervised deep learning-based 3D scene understanding methods can overcome the difficulty of acquiring the densely labeled ground-truth and have made a lot of advances. However, occlusions and moving objects are still some of the major…

计算机视觉与模式识别 · 计算机科学 2021-10-25 Jiaojiao Fang , Guizhong Liu

Reconstructing clothed humans from a single image is a fundamental task in computer vision with wide-ranging applications. Although existing monocular clothed human reconstruction solutions have shown promising results, they often rely on…

计算机视觉与模式识别 · 计算机科学 2025-10-20 Arindam Dutta , Meng Zheng , Zhongpai Gao , Benjamin Planche , Anwesha Choudhuri , Terrence Chen , Amit K. Roy-Chowdhury , Ziyan Wu

With the availability of egocentric 3D hand-object interaction datasets, there is increasing interest in developing unified models for hand-object pose estimation and action recognition. However, existing methods still struggle to recognise…

计算机视觉与模式识别 · 计算机科学 2025-01-14 Tze Ho Elden Tse , Runyang Feng , Linfang Zheng , Jiho Park , Yixing Gao , Jihie Kim , Ales Leonardis , Hyung Jin Chang

Despite their irresistible success, deep learning algorithms still heavily rely on annotated data. On the other hand, unsupervised settings pose many challenges, especially about determining the right inductive bias in diverse scenarios.…

计算机视觉与模式识别 · 计算机科学 2021-03-11 Beril Besbinar , Pascal Frossard

Amodal perception requires inferring the full shape of an object that is partially occluded. This task is particularly challenging on two levels: (1) it requires more information than what is contained in the instant retina or imaging…

计算机视觉与模式识别 · 计算机科学 2022-10-25 Jian Yao , Yuxin Hong , Chiyu Wang , Tianjun Xiao , Tong He , Francesco Locatello , David Wipf , Yanwei Fu , Zheng Zhang

Accurately modeling detailed interactions between human/hand and object is an appealing yet challenging task. Current multi-view capture systems are only capable of reconstructing multiple subjects into a single, unified mesh, which fails…

计算机视觉与模式识别 · 计算机科学 2024-03-22 Jiajun Zhang , Yuxiang Zhang , Hongwen Zhang , Xiao Zhou , Boyao Zhou , Ruizhi Shao , Zonghai Hu , Yebin Liu

Task-oriented object grasping and rearrangement are critical skills for robots to accomplish different real-world manipulation tasks. However, they remain challenging due to partial observations of the objects and shape variations in…

机器人学 · 计算机科学 2026-03-06 Yichen Cai , Jianfeng Gao , Christoph Pohl , Tamim Asfour

We propose a robust and accurate method for reconstructing 3D hand mesh from monocular images. This is a very challenging problem, as hands are often severely occluded by objects. Previous works often have disregarded 2D hand pose…

计算机视觉与模式识别 · 计算机科学 2024-03-14 Shuaibing Wang , Shunli Wang , Dingkang Yang , Mingcheng Li , Ziyun Qian , Liuzhen Su , Lihua Zhang

Monocular 3D reconstruction of articulated object categories is challenging due to the lack of training data and the inherent ill-posedness of the problem. In this work we use video self-supervision, forcing the consistency of consecutive…

计算机视觉与模式识别 · 计算机科学 2021-04-28 Filippos Kokkinos , Iasonas Kokkinos

Ego-motion estimation is vital for drones when flying in GPS-denied environments. Vision-based methods struggle when flight speed increases and close-by objects lead to difficult visual conditions with considerable motion blur and large…

机器人学 · 计算机科学 2025-05-01 Stavrow A. Bahnam , Christophe De Wagter , Guido C. H. E. de Croon

Generating realistic hand-object interactions (HOI) videos is a significant challenge due to the difficulty of modeling physical constraints (e.g., contact and occlusion between hands and manipulated objects). Current methods utilize HOI…

计算机视觉与模式识别 · 计算机科学 2025-12-02 Haodong Yan , Hang Yu , Zhide Zhong , Weilin Yuan , Xin Gong , Zehang Luo , Chengxi Heyu , Junfeng Li , Wenxuan Song , Shunbo Zhou , Haoang Li

The challenge of dynamic view synthesis from dynamic monocular videos, i.e., synthesizing novel views for free viewpoints given a monocular video of a dynamic scene captured by a moving camera, mainly lies in accurately modeling the…

计算机视觉与模式识别 · 计算机科学 2024-08-22 Meng You , Junhui Hou

We introduce a novel framework for reconstructing dynamic human-object interactions from monocular video that overcomes challenges associated with occlusions and temporal inconsistencies. Traditional 3D reconstruction methods typically…

计算机视觉与模式识别 · 计算机科学 2025-09-16 Hyungjun Doh , Dong In Lee , Seunggeun Chi , Pin-Hao Huang , Kwonjoon Lee , Sangpil Kim , Karthik Ramani

Reconstructing 3D models of dynamic, real-world objects with high-fidelity textures from monocular frame sequences has been a challenging problem in recent years. This difficulty stems from factors such as shadows, indirect illumination,…

计算机视觉与模式识别 · 计算机科学 2025-05-12 Alakh Aggarwal , Ningna Wang , Xiaohu Guo

Unsupervised multi-object segmentation has shown impressive results on images by utilizing powerful semantics learned from self-supervised pretraining. An additional modality such as depth or motion is often used to facilitate the…

计算机视觉与模式识别 · 计算机科学 2023-10-12 Görkay Aydemir , Weidi Xie , Fatma Güney