English
Related papers

Related papers: Endo3R: Unified Online Reconstruction from Dynamic…

200 papers

Traditional high-quality 3D scanning and reconstruction typically relies on human labor to plan the scanning procedure. With the rapid development of embodied systems such as drones and robots, there is a growing demand of performing…

Computer Vision and Pattern Recognition · Computer Science 2025-12-05 Chentao Shen , Sizhe Zheng , Bingqian Wu , Yaohua Feng , Yuanchen Fei , Mingyu Mei , Hanwen Jiang , Xiangru Huang

Accurate monocular metric depth estimation (MMDE) is crucial to solving downstream tasks in 3D perception and modeling. However, the remarkable accuracy of recent MMDE methods is confined to their training domains. These methods fail to…

Computer Vision and Pattern Recognition · Computer Science 2025-12-19 Luigi Piccinelli , Christos Sakaridis , Yung-Hsu Yang , Mattia Segu , Siyuan Li , Wim Abbeloos , Luc Van Gool

In this paper, we tackle the problem of estimating the depth of a scene from a monocular video sequence. In particular, we handle challenging scenarios, such as non-translational camera motion and dynamic scenes, where traditional structure…

Computer Vision and Pattern Recognition · Computer Science 2015-11-20 Miaomiao Liu , Mathieu Salzmann , Xuming He

We present an approach that learns to synthesize high-quality, novel views of 3D objects or scenes, while providing fine-grained and precise control over the 6-DOF viewpoint. The approach is self-supervised and only requires 2D images and…

Computer Vision and Pattern Recognition · Computer Science 2019-09-10 Xu Chen , Jie Song , Otmar Hilliges

Monocular visual localization plays a pivotal role in advanced driver assistance systems and autonomous driving by estimating a vehicle's ego-motion from a single pinhole camera. Nevertheless, conventional monocular visual odometry…

Computer Vision and Pattern Recognition · Computer Science 2025-03-07 Hui Zhang , Zhiyang Wu , Qianqian Shangguan , Kang An

One image to editable dynamic 3D model and video generation is novel direction and change in the research area of single image to 3D representation or 3D reconstruction of image. Gaussian Splatting has demonstrated its advantages in…

Computer Vision and Pattern Recognition · Computer Science 2024-05-13 Jinwei Lin

Recovering the 3D geometry of a scene from a sparse set of uncalibrated images is a long-standing problem in computer vision. While recent learning-based approaches such as DUSt3R and MASt3R have demonstrated impressive results by directly…

Computer Vision and Pattern Recognition · Computer Science 2025-08-25 Sara Rojas , Matthieu Armando , Bernard Ghamen , Philippe Weinzaepfel , Vincent Leroy , Gregory Rogez

UAVs have become an essential photogrammetric measurement as they are affordable, easily accessible and versatile. Aerial images captured from UAVs have applications in small and large scale texture mapping, 3D modelling, object detection…

Computer Vision and Pattern Recognition · Computer Science 2020-12-22 Logambal Madhuanand , Francesco Nex , Michael Ying Yang

We introduce POP3D, a novel framework that creates a full $360^\circ$-view 3D model from a single image. POP3D resolves two prominent issues that limit the single-view reconstruction. Firstly, POP3D offers substantial generalizability to…

Computer Vision and Pattern Recognition · Computer Science 2023-09-20 Nuri Ryu , Minsu Gong , Geonung Kim , Joo-Haeng Lee , Sunghyun Cho

Creating a photorealistic scene and human reconstruction from a single monocular in-the-wild video figures prominently in the perception of a human-centric 3D world. Recent neural rendering advances have enabled holistic human-scene…

Computer Vision and Pattern Recognition · Computer Science 2025-04-21 Zetong Zhang , Manuel Kaufmann , Lixin Xue , Jie Song , Martin R. Oswald

Existing methods for reconstructing objects and humans from a monocular image suffer from severe mesh collisions and performance limitations for interacting occluding objects. This paper introduces a method to obtain a globally consistent…

Computer Vision and Pattern Recognition · Computer Science 2024-08-16 Sarthak Batra , Partha P. Chakrabarti , Simon Hadfield , Armin Mustafa

3D reconstruction of deformable (or non-rigid) scenes from a set of monocular 2D image observations is a long-standing and actively researched area of computer vision and graphics. It is an ill-posed inverse problem, since -- without…

Computer Vision and Pattern Recognition · Computer Science 2023-03-28 Edith Tretschk , Navami Kairanda , Mallikarjun B R , Rishabh Dabral , Adam Kortylewski , Bernhard Egger , Marc Habermann , Pascal Fua , Christian Theobalt , Vladislav Golyanik

In this paper, we focus on online zero-shot monocular 3D instance segmentation, a novel practical setting where existing approaches fail to perform because they rely on posed RGB-D sequences. To overcome this limitation, we leverage CUT3R,…

Computer Vision and Pattern Recognition · Computer Science 2026-04-22 Zhipeng Du , Duolikun Danier , Jan Eric Lenssen , Hakan Bilen

In this paper, we propose MonoRec, a semi-supervised monocular dense reconstruction architecture that predicts depth maps from a single moving camera in dynamic environments. MonoRec is based on a multi-view stereo setting which encodes the…

Computer Vision and Pattern Recognition · Computer Science 2022-09-22 Felix Wimbauer , Nan Yang , Lukas von Stumberg , Niclas Zeller , Daniel Cremers

The rapid advancement of Large Multimodal Models (LMMs) for 2D images and videos has motivated extending these models to understand 3D scenes, aiming for human-like visual-spatial intelligence. Nevertheless, achieving deep spatial…

We present UniScale, a unified, scale-aware multi-view 3D reconstruction framework for robotic applications that flexibly integrates geometric priors through a modular, semantically informed design. In vision-based robotic navigation, the…

Computer Vision and Pattern Recognition · Computer Science 2026-02-27 Mohammad Mahdavian , Gordon Tan , Binbin Xu , Yuan Ren , Dongfeng Bai , Bingbing Liu

We present a novel method to reconstruct 3D scenes from images by leveraging deep dense monocular SLAM and fast uncertainty propagation. The proposed approach is able to 3D reconstruct scenes densely, accurately, and in real-time while…

Computer Vision and Pattern Recognition · Computer Science 2022-10-18 Antoni Rosinol , John J. Leonard , Luca Carlone

We propose a novel framework for scene decomposition and static background reconstruction from everyday videos. By integrating the trained motion masks and modeling the static scene as Gaussian splats with dynamics-aware optimization, our…

Computer Vision and Pattern Recognition · Computer Science 2024-12-30 Kai Xu , Tze Ho Elden Tse , Jizong Peng , Angela Yao

Recent feed-forward reconstruction models like VGGT and $\pi^3$ achieve impressive reconstruction quality but cannot process streaming videos due to quadratic memory complexity, limiting their practical deployment. While existing streaming…

Computer Vision and Pattern Recognition · Computer Science 2026-03-31 Tianye Ding , Yiming Xie , Yiqing Liang , Moitreya Chatterjee , Pedro Miraldo , Huaizu Jiang

Monocular depth prediction plays a crucial role in understanding 3D scene geometry. Although recent methods have achieved impressive progress in terms of evaluation metrics such as the pixel-wise relative error, most methods neglect the…

Computer Vision and Pattern Recognition · Computer Science 2021-06-29 Wei Yin , Yifan Liu , Chunhua Shen