中文
相关论文

相关论文: PAD3R: Pose-Aware Dynamic 3D Reconstruction from C…

200 篇论文

This paper describes how to obtain accurate 3D body models and texture of arbitrary people from a single, monocular video in which a person is moving. Based on a parametric body model, we present a robust processing pipeline achieving 3D…

计算机视觉与模式识别 · 计算机科学 2018-04-17 Thiemo Alldieck , Marcus Magnor , Weipeng Xu , Christian Theobalt , Gerard Pons-Moll

Reconstructing dynamic 3D scenes from monocular videos requires simultaneously capturing high-frequency appearance details and temporally continuous motion. Existing methods using single Gaussian primitives are limited by their low-pass…

计算机视觉与模式识别 · 计算机科学 2026-01-05 Jiewen Chan , Zhenjun Zhao , Yu-Lun Liu

The recent state of the art on monocular 3D face reconstruction from image data has made some impressive advancements, thanks to the advent of Deep Learning. However, it has mostly focused on input coming from a single RGB image,…

We introduce a novel framework for reconstructing dynamic human-object interactions from monocular video that overcomes challenges associated with occlusions and temporal inconsistencies. Traditional 3D reconstruction methods typically…

计算机视觉与模式识别 · 计算机科学 2025-09-16 Hyungjun Doh , Dong In Lee , Seunggeun Chi , Pin-Hao Huang , Kwonjoon Lee , Sangpil Kim , Karthik Ramani

Learning to understand dynamic 3D scenes from imagery is crucial for applications ranging from robotics to scene reconstruction. Yet, unlike other problems where large-scale supervised training has enabled rapid progress, directly…

计算机视觉与模式识别 · 计算机科学 2025-05-01 Linyi Jin , Richard Tucker , Zhengqi Li , David Fouhey , Noah Snavely , Aleksander Holynski

In this paper, we introduce Splatt3R, a pose-free, feed-forward method for in-the-wild 3D reconstruction and novel view synthesis from stereo pairs. Given uncalibrated natural images, Splatt3R can predict 3D Gaussian Splats without…

计算机视觉与模式识别 · 计算机科学 2024-08-29 Brandon Smart , Chuanxia Zheng , Iro Laina , Victor Adrian Prisacariu

Reconstructing hand-held objects in 3D from monocular images remains a significant challenge in computer vision. Most existing approaches rely on implicit 3D representations, which produce overly smooth reconstructions and are…

计算机视觉与模式识别 · 计算机科学 2025-07-22 Zerui Chen , Rolandos Alexandros Potamias , Shizhe Chen , Cordelia Schmid

We present a novel approach to interactive 3D object perception for robots. Unlike previous perception algorithms that rely on known object models or a large amount of annotated training data, we propose a poking-based approach that…

机器人学 · 计算机科学 2023-02-28 Linghao Chen , Yunzhou Song , Hujun Bao , Xiaowei Zhou

We present a self-supervised learning-based pipeline for dense 3D reconstruction from full-length monocular endoscopic videos without a priori modeling of anatomy or shading. Our method only relies on unlabeled monocular endoscopic videos…

计算机视觉与模式识别 · 计算机科学 2019-09-10 Xingtong Liu , Ayushi Sinha , Masaru Ishii , Gregory D. Hager , Russell H. Taylor , Mathias Unberath

Understanding dynamic scenes from casual videos is critical for scalable robot learning, yet four-dimensional (4D) reconstruction under strictly monocular settings remains highly ill-posed. To address this challenge, our key insight is that…

计算机视觉与模式识别 · 计算机科学 2026-02-17 Can Li , Jie Gu , Jingmin Chen , Fangzhou Qiu , Lei Sun

Estimating accurate camera poses, 3D scene geometry, and object motion from in-the-wild videos is a long-standing challenge for classical structure from motion pipelines due to the presence of dynamic objects. Recent learning-based methods…

计算机视觉与模式识别 · 计算机科学 2025-12-08 Zhuoyuan Wu , Xurui Yang , Jiahui Huang , Yue Wang , Jun Gao

We present an approach for 3D global human mesh recovery from monocular videos recorded with dynamic cameras. Our approach is robust to severe and long-term occlusions and tracks human bodies even when they go outside the camera's field of…

计算机视觉与模式识别 · 计算机科学 2022-03-31 Ye Yuan , Umar Iqbal , Pavlo Molchanov , Kris Kitani , Jan Kautz

The challenge of dynamic view synthesis from dynamic monocular videos, i.e., synthesizing novel views for free viewpoints given a monocular video of a dynamic scene captured by a moving camera, mainly lies in accurately modeling the…

计算机视觉与模式识别 · 计算机科学 2024-08-22 Meng You , Junhui Hou

Dense 3D tracking from monocular video is fundamental to dynamic scene understanding. While recent 3D foundation models provide reliable per-frame geometry, recovering object motion in this geometry remains challenging and benefits from…

计算机视觉与模式识别 · 计算机科学 2026-05-14 Jisu Nam , Jahyeok Koo , Soowon Son , Jaewoo Jung , Honggyu An , Junhwa Hur , Seungryong Kim

Understanding 3D motion from videos presents inherent challenges due to the diverse types of movement, ranging from rigid and deformable objects to articulated structures. To overcome this, we propose Liv3Stroke, a novel approach for…

计算机视觉与模式识别 · 计算机科学 2025-03-28 Jaeah Lee , Changwoon Choi , Young Min Kim , Jaesik Park

We present PlanarRecon -- a novel framework for globally coherent detection and reconstruction of 3D planes from a posed monocular video. Unlike previous works that detect planes in 2D from a single image, PlanarRecon incrementally detects…

计算机视觉与模式识别 · 计算机科学 2022-06-16 Yiming Xie , Matheus Gadelha , Fengting Yang , Xiaowei Zhou , Huaizu Jiang

It has long been challenging to recover the underlying dynamic 3D scene representations from a monocular RGB video. Existing works formulate this problem into finding a single most plausible solution by adding various constraints such as…

计算机视觉与模式识别 · 计算机科学 2024-07-09 Ziyang Song , Jinxi Li , Bo Yang

Recent advancements in multi-view scene reconstruction have been significant, yet existing methods face limitations when processing streams of input images. These methods either rely on time-consuming offline optimization or are restricted…

计算机视觉与模式识别 · 计算机科学 2025-07-25 Zhuoguang Chen , Minghui Qin , Tianyuan Yuan , Zhe Liu , Hang Zhao

This paper proposes a new approach for monocular dense 3D reconstruction of a complex dynamic scene from two perspective frames. By applying superpixel over-segmentation to the image, we model a generically dynamic (hence non-rigid) scene…

计算机视觉与模式识别 · 计算机科学 2017-12-21 Suryansh Kumar , Yuchao Dai , Hongdong Li

Narrated instructional videos often show and describe manipulations of similar objects, e.g., repairing a particular model of a car or laptop. In this work we aim to reconstruct such objects and to localize associated narrations in 3D.…

计算机视觉与模式识别 · 计算机科学 2021-09-13 Dimitri Zhukov , Ignacio Rocco , Ivan Laptev , Josef Sivic , Johannes L. Schönberger , Bugra Tekin , Marc Pollefeys