English
Related papers

Related papers: ForeHOI: Feed-forward 3D Object Reconstruction fro…

200 papers

Monocular 3D scene reconstruction has recently seen significant progress. Powered by the modern neural architectures and large-scale data, recent methods achieve high performance in depth estimation from a single image. Meanwhile,…

Computer Vision and Pattern Recognition · Computer Science 2026-03-18 Junaid Ahmed Ansari , Ran Ding , Fabio Pizzati , Ivan Laptev

Dynamic driving scene reconstruction is critical for autonomous driving simulation and closed-loop learning. While recent feed-forward methods have shown promise for 3D reconstruction, they struggle with long-range driving sequences due to…

Computer Vision and Pattern Recognition · Computer Science 2026-02-25 Kaiyuan Tan , Yingying Shen , Mingfei Tu , Haohui Zhu , Bing Wang , Guang Chen , Hangjun Ye , Haiyang Sun

3D human reconstruction and animation are long-standing topics in computer graphics and vision. However, existing methods typically rely on sophisticated dense-view capture and/or time-consuming per-subject optimization procedures. To…

Graphics · Computer Science 2025-06-04 Zhiyuan Yu , Zhe Li , Hujun Bao , Can Yang , Xiaowei Zhou

Generating physically realistic humanoid-object interactions (HOI) is a fundamental challenge in robotics. Existing HOI generation approaches, such as diffusion-based models, often suffer from artifacts such as implausible contacts,…

Robotics · Computer Science 2025-08-21 Yuhang Lin , Yijia Xie , Jiahong Xie , Yuehao Huang , Ruoyu Wang , Jiajun Lv , Yukai Ma , Xingxing Zuo

There has been much recent interest in deep learning methods for monocular image based object pose estimation. While object pose estimation is an important problem for autonomous robot interaction with the physical world, and the…

Computer Vision and Pattern Recognition · Computer Science 2020-03-02 Gideon Billings , Matthew Johnson-Roberson

Real-time 3D hand forecasting is a critical component for fluid human-computer interaction in applications like AR and assistive robotics. However, existing methods are ill-suited for these scenarios, as they typically require offline…

Computer Vision and Pattern Recognition · Computer Science 2026-05-18 Ruicong Liu , Yifei Huang , Liangyang Ouyang , Caixin Kang , Yoichi Sato

Humans interact with objects all the time. Enabling a humanoid to learn human-object interaction (HOI) is a key step for future smart animation and intelligent robotics systems. However, recent progress in physics-based HOI requires…

Computer Vision and Pattern Recognition · Computer Science 2023-12-08 Yinhuai Wang , Jing Lin , Ailing Zeng , Zhengyi Luo , Jian Zhang , Lei Zhang

Enabling humanoid robots to clean rooms has long been a pursued dream within humanoid research communities. However, many tasks require multi-humanoid collaboration, such as carrying large and heavy furniture together. Given the scarcity of…

Robotics · Computer Science 2024-10-31 Jiawei Gao , Ziqin Wang , Zeqi Xiao , Jingbo Wang , Tai Wang , Jinkun Cao , Xiaolin Hu , Si Liu , Jifeng Dai , Jiangmiao Pang

Hand-object interaction(HOI) is the fundamental link between human and environment, yet its dexterous and complex pose significantly challenges for gesture control. Despite significant advances in AI and robotics, enabling machines to…

Robotics · Computer Science 2025-07-11 Yongqi Tian , Xueyu Sun , Haoyuan He , Linji Hao , Ning Ding , Caigui Jiang

We introduce a novel camera model for monocular 3D Morphable Model (3DMM) regression methods that effectively captures the perspective distortion effect commonly seen in close-up facial images. Fitting 3D morphable models to video is a key…

Computer Vision and Pattern Recognition · Computer Science 2026-03-06 Toby Chong , Ryota Nakajima

Monocular depth estimation has been actively studied in fields such as robot vision, autonomous driving, and 3D scene understanding. Given a sequence of color images, unsupervised learning methods based on the framework of…

Computer Vision and Pattern Recognition · Computer Science 2022-12-06 Songlin Wei , Guodong Chen , Wenzheng Chi , Zhenhua Wang , Lining Sun

While 3D hand reconstruction from monocular images has made significant progress, generating accurate and temporally coherent motion estimates from videos remains challenging, particularly during hand-object interactions. In this paper, we…

Computer Vision and Pattern Recognition · Computer Science 2025-08-05 Yufei Zhang , Zijun Cui , Jeffrey O. Kephart , Qiang Ji

We present the first method for real-time full body capture that estimates shape and motion of body and hands together with a dynamic 3D face model from a single color image. Our approach uses a new neural network architecture that exploits…

Computer Vision and Pattern Recognition · Computer Science 2021-04-16 Yuxiao Zhou , Marc Habermann , Ikhsanul Habibie , Ayush Tewari , Christian Theobalt , Feng Xu

We present a real-time deep learning framework for video-based facial performance capture -- the dense 3D tracking of an actor's face given a monocular video. Our pipeline begins with accurately capturing a subject using a high-end…

Computer Vision and Pattern Recognition · Computer Science 2017-06-05 Samuli Laine , Tero Karras , Timo Aila , Antti Herva , Shunsuke Saito , Ronald Yu , Hao Li , Jaakko Lehtinen

Hand-Object Interaction (HOI) generation plays a critical role in advancing applications across animation and robotics. Current video-based methods are predominantly single-view, which impedes comprehensive 3D geometry perception and often…

Computer Vision and Pattern Recognition · Computer Science 2026-03-09 Lingwei Dang , Zonghan Li , Juntong Li , Hongwen Zhang , Liang An , Yebin Liu , Qingyao Wu

Temporal 3D human pose estimation from monocular videos is a challenging task in human-centered computer vision due to the depth ambiguity of 2D-to-3D lifting. To improve accuracy and address occlusion issues, inertial sensor has been…

Computer Vision and Pattern Recognition · Computer Science 2024-04-30 Yiming Bao , Xu Zhao , Dahong Qian

Gaze plays a crucial role in revealing human attention and intention, particularly in hand-object interaction scenarios, where it guides and synchronizes complex tasks that require precise coordination between the brain, hand, and object.…

Computer Vision and Pattern Recognition · Computer Science 2026-03-18 Jie Tian , Ran Ji , Lingxiao Yang , Suting Ni , Yuexin Ma , Lan Xu , Jingyi Yu , Ye Shi , Jingya Wang

Reconstructing complete and animatable 3D human avatars from monocular videos remains challenging, particularly under severe occlusions. While 3D Gaussian Splatting has enabled photorealistic human rendering, existing methods struggle with…

Computer Vision and Pattern Recognition · Computer Science 2026-01-06 Jinlong Fan , Shanshan Zhao , Liang Zheng , Jing Zhang , Yuxiang Yang , Mingming Gong

We propose a method for in-hand 3D scanning of an unknown object with a monocular camera. Our method relies on a neural implicit surface representation that captures both the geometry and the appearance of the object, however, by contrast…

Computer Vision and Pattern Recognition · Computer Science 2023-06-23 Shreyas Hampali , Tomas Hodan , Luan Tran , Lingni Ma , Cem Keskin , Vincent Lepetit

Hand-object 3D reconstruction has become increasingly important for applications in human-robot interaction and immersive AR/VR experiences. A common approach for object-agnostic hand-object reconstruction from RGB sequences involves a…

Computer Vision and Pattern Recognition · Computer Science 2025-08-26 Anilkumar Swamy , Vincent Leroy , Philippe Weinzaepfel , Jean-Sébastien Franco , Grégory Rogez