English
Related papers

Related papers: PAS3R: Pose-Adaptive Streaming 3D Reconstruction f…

200 papers

Feedforward geometric foundation models achieve strong short-window reconstruction, yet scaling them to minutes-long videos is bottlenecked by quadratic attention complexity or limited effective memory in recurrent designs. We present LoGeR…

Computer Vision and Pattern Recognition · Computer Science 2026-04-28 Junyi Zhang , Charles Herrmann , Junhwa Hur , Chen Sun , Ming-Hsuan Yang , Forrester Cole , Trevor Darrell , Deqing Sun

Reconstructing scenes and tracking motion are two sides of the same coin. Tracking points allow for geometric reconstruction [14], while geometric reconstruction of (dynamic) scenes allows for 3D tracking of points over time [24, 39]. The…

Computer Vision and Pattern Recognition · Computer Science 2025-03-13 Jenny Seidenschwarz , Qunjie Zhou , Bardienus Duisterhof , Deva Ramanan , Laura Leal-Taixé

Real-time, high-quality, 3D scanning of large-scale scenes is key to mixed reality and robotic applications. However, scalability brings challenges of drift in pose estimation, introducing significant errors in the accumulated model.…

Graphics · Computer Science 2017-02-09 Angela Dai , Matthias Nießner , Michael Zollhöfer , Shahram Izadi , Christian Theobalt

Current feed-forward 3D/4D reconstruction systems rely on dense geometry and pose supervision -- expensive to obtain at scale and particularly scarce for dynamic real-world scenes. We present Flow3r, a framework that augments visual…

Computer Vision and Pattern Recognition · Computer Science 2026-02-24 Zhongxiao Cong , Qitao Zhao , Minsik Jeon , Shubham Tulsiani

Modern Recurrent Neural Networks have become a competitive architecture for 3D reconstruction due to their linear-time complexity. However, their performance degrades significantly when applied beyond the training context length, revealing…

Computer Vision and Pattern Recognition · Computer Science 2026-03-04 Xingyu Chen , Yue Chen , Yuliang Xiu , Andreas Geiger , Anpei Chen

Reconstructing dynamic articulated objects from a singular monocular video is challenging, requiring joint estimation of shape, motion, and camera parameters from limited views. Current methods typically demand extensive computational…

Computer Vision and Pattern Recognition · Computer Science 2024-05-22 Hao Zhang , Fang Li , Samyak Rawlekar , Narendra Ahuja

The creation of detailed 3D models is relevant for a wide range of applications such as navigation in three-dimensional space, construction planning or disaster assessment. However, the complex processing and long execution time for…

Computer Vision and Pattern Recognition · Computer Science 2021-10-25 Max Hermann , Thomas Pollok , Daniel Brommer , Dominic Zahn

LongSplat addresses critical challenges in novel view synthesis (NVS) from casually captured long videos characterized by irregular camera motion, unknown camera poses, and expansive scenes. Current methods often suffer from pose drift,…

Computer Vision and Pattern Recognition · Computer Science 2025-08-20 Chin-Yang Lin , Cheng Sun , Fu-En Yang , Min-Hung Chen , Yen-Yu Lin , Yu-Lun Liu

Recent advances in generalizable 3D Gaussian Splatting (3DGS) have enabled rapid 3D scene reconstruction within seconds, eliminating the need for per-scene optimization. However, existing methods primarily follow an offline reconstruction…

Computer Vision and Pattern Recognition · Computer Science 2026-03-04 Chong Xia , Fangfu Liu , Yule Wang , Yize Pang , Yueqi Duan

Reconstructing an accurate 3D object model from a few image observations remains a challenging problem in computer vision. State-of-the-art approaches typically assume accurate camera poses as input, which could be difficult to obtain in…

Computer Vision and Pattern Recognition · Computer Science 2022-05-17 Zhenpei Yang , Zhile Ren , Miguel Angel Bautista , Zaiwei Zhang , Qi Shan , Qixing Huang

Photorealistic 3-D reconstruction from monocular video collapses in large-scale scenes when depth, pose, and radiance are solved in isolation: scale-ambiguous depth yields ghost geometry, long-horizon pose drift corrupts alignment, and a…

Computer Vision and Pattern Recognition · Computer Science 2026-02-02 Shahram Najam Syed , Yitian Hu , Yuchao Yao

We present ReFlow, a unified framework for monocular dynamic scene reconstruction that learns 3D motion in a novel self-correction manner from raw video. Existing methods often suffer from incomplete scene initialization for dynamic…

Computer Vision and Pattern Recognition · Computer Science 2026-04-03 Yanzhe Liang , Ruijie Zhu , Hanzhi Chang , Zhuoyuan Li , Jiahao Lu , Tianzhu Zhang

We introduce UniCon3R, a unified feed-forward framework for online human-scene 4D reconstruction from monocular video. Current feed-forward human-scene reconstruction methods suffer from artifacts, where bodies float above the ground or…

Computer Vision and Pattern Recognition · Computer Science 2026-05-12 Tanuj Sur , Shashank Tripathi , Nikos Athanasiou , Ha Linh Nguyen , Kai Xu , Michael J. Black , Angela Yao

Accurate depth and camera pose estimation is essential for achieving high-quality 3D visualisations in robotic-assisted surgery. Despite recent advancements in foundation model adaptation to monocular depth estimation of endoscopic scenes…

Computer Vision and Pattern Recognition · Computer Science 2025-08-21 Mona Sheikh Zeinoddin , Mobarak I. Hoque , Zafer Tandogdu , Greg Shaw , Matthew J. Clarkson , Evangelos Mazomenos , Danail Stoyanov

Online test-time adaptation for 3D human pose estimation is used for video streams that differ from training data. Ground truth 2D poses are used for adaptation, but only estimated 2D poses are available in practice. This paper addresses…

Computer Vision and Pattern Recognition · Computer Science 2025-03-17 Qiuxia Lin , Kerui Gu , Linlin Yang , Angela Yao

We present an algorithm for estimating consistent dense depth maps and camera poses from a monocular video. We integrate a learning-based depth prior, in the form of a convolutional neural network trained for single-image depth estimation,…

Computer Vision and Pattern Recognition · Computer Science 2021-06-23 Johannes Kopf , Xuejian Rong , Jia-Bin Huang

Traditional SLAM systems, which rely on bundle adjustment, struggle with highly dynamic scenes commonly found in casual videos. Such videos entangle the motion of dynamic elements, undermining the assumption of static environments required…

Computer Vision and Pattern Recognition · Computer Science 2025-11-07 Weirong Chen , Ganlin Zhang , Felix Wimbauer , Rui Wang , Nikita Araslanov , Andrea Vedaldi , Daniel Cremers

The precise reconstruction of 3D objects from a single RGB image in complex scenes presents a critical challenge in virtual reality, autonomous driving, and robotics. Existing neural implicit 3D representation methods face significant…

Computer Vision and Pattern Recognition · Computer Science 2024-11-21 Luoxi Zhang , Pragyan Shrestha , Yu Zhou , Chun Xie , Itaru Kitahara

Advancements in 3D scene reconstruction have transformed 2D images from the real world into 3D models, producing realistic 3D results from hundreds of input photos. Despite great success in dense-view reconstruction scenarios, rendering a…

Computer Vision and Pattern Recognition · Computer Science 2025-06-26 Fangfu Liu , Wenqiang Sun , Hanyang Wang , Yikai Wang , Haowen Sun , Junliang Ye , Jun Zhang , Yueqi Duan

Recent advances in 2D-to-3D perception have enabled the recovery of 3D scene semantics from unposed images. However, prevailing methods often suffer from limited generalization, reliance on per-scene optimization, and semantic…

Computer Vision and Pattern Recognition · Computer Science 2026-03-27 Jie Hu , Shizun Wang , Xinchao Wang