English
Related papers

Related papers: Endo3R: Unified Online Reconstruction from Dynamic…

200 papers

Monocular depth reconstruction of complex and dynamic scenes is a highly challenging problem. While for rigid scenes learning-based methods have been offering promising results even in unsupervised cases, there exists little to no…

Computer Vision and Pattern Recognition · Computer Science 2021-10-29 Ayça Takmaz , Danda Pani Paudel , Thomas Probst , Ajad Chhatkuli , Martin R. Oswald , Luc Van Gool

We present WinT3R, a feed-forward reconstruction model capable of online prediction of precise camera poses and high-quality point maps. Previous methods suffer from a trade-off between reconstruction quality and real-time performance. To…

Computer Vision and Pattern Recognition · Computer Science 2025-09-08 Zizun Li , Jianjun Zhou , Yifan Wang , Haoyu Guo , Wenzheng Chang , Yang Zhou , Haoyi Zhu , Junyi Chen , Chunhua Shen , Tong He

The application of monocular dense Simultaneous Localization and Mapping (SLAM) is often hindered by high latency, large GPU memory consumption, and reliance on camera calibration. To relax this constraint, we propose EC3R-SLAM, a novel…

Robotics · Computer Science 2025-10-03 Lingxiang Hu , Naima Ait Oufroukh , Fabien Bonardi , Raymond Ghandour

Depth estimation is a cornerstone of 3D reconstruction and plays a vital role in minimally invasive endoscopic surgeries. However, most current depth estimation networks rely on traditional convolutional neural networks, which are limited…

Computer Vision and Pattern Recognition · Computer Science 2025-07-16 Bojian Li , Bo Liu , Xinning Yao , Jinghua Yue , Fugen Zhou

Reconstructing soft tissues from stereo endoscope videos is an essential prerequisite for many medical applications. Previous methods struggle to produce high-quality geometry and appearance due to their inadequate representations of 3D…

Computer Vision and Pattern Recognition · Computer Science 2023-09-06 Ruyi Zha , Xuelian Cheng , Hongdong Li , Mehrtash Harandi , Zongyuan Ge

We present Edit3r, a feed-forward framework that reconstructs and edits 3D scenes in a single pass from unposed, view-inconsistent, instruction-edited images. Unlike prior methods requiring per-scene optimization, Edit3r directly predicts…

Computer Vision and Pattern Recognition · Computer Science 2026-01-01 Jiageng Liu , Weijie Lyu , Xueting Li , Yejie Guo , Ming-Hsuan Yang

Reconstruction of the soft tissues in robotic surgery from endoscopic stereo videos is important for many applications such as intra-operative navigation and image-guided robotic surgery automation. Previous works on this task mainly rely…

Computer Vision and Pattern Recognition · Computer Science 2022-07-01 Yuehao Wang , Yonghao Long , Siu Hin Fan , Qi Dou

Learning to understand dynamic 3D scenes from imagery is crucial for applications ranging from robotics to scene reconstruction. Yet, unlike other problems where large-scale supervised training has enabled rapid progress, directly…

Computer Vision and Pattern Recognition · Computer Science 2025-05-01 Linyi Jin , Richard Tucker , Zhengqi Li , David Fouhey , Noah Snavely , Aleksander Holynski

Recovering temporally consistent 3D human body pose, shape and motion from a monocular video is a challenging task due to (self-)occlusions, poor lighting conditions, complex articulated body poses, depth ambiguity, and limited availability…

Computer Vision and Pattern Recognition · Computer Science 2023-11-21 Sushovan Chanda , Amogh Tiwari , Lokender Tiwari , Brojeshwar Bhowmick , Avinash Sharma , Hrishav Barua

We present a novel method to learn temporally consistent 3D reconstruction of clothed people from a monocular video. Recent methods for 3D human reconstruction from monocular video using volumetric, implicit or parametric human shape…

Computer Vision and Pattern Recognition · Computer Science 2021-04-20 Akin Caliskan , Armin Mustafa , Adrian Hilton

We present a novel framework named NeuralRecon for real-time 3D scene reconstruction from a monocular video. Unlike previous methods that estimate single-view depth maps separately on each key-frame and fuse them later, we propose to…

Computer Vision and Pattern Recognition · Computer Science 2021-04-02 Jiaming Sun , Yiming Xie , Linghao Chen , Xiaowei Zhou , Hujun Bao

In robot-assisted minimally invasive surgery, accurate 3D reconstruction from endoscopic video is vital for downstream tasks and improved outcomes. However, endoscopic scenarios present unique challenges, including photometric…

Computer Vision and Pattern Recognition · Computer Science 2025-10-28 Taoyu Wu , Yiyi Miao , Jiaxin Guo , Ziyan Chen , Sihang Zhao , Zhuoxiao Li , Zhe Tang , Baoru Huang , Limin Yu

Panoptic 3D reconstruction from a monocular video is a fundamental perceptual task in robotic scene understanding. However, existing efforts suffer from inefficiency in terms of inference speed and accuracy, limiting their practical…

Computer Vision and Pattern Recognition · Computer Science 2024-09-23 Zhen Zhou , Yunkai Ma , Junfeng Fan , Shaolin Zhang , Fengshui Jing , Min Tan

Multi-view 3D reconstruction remains a core challenge in computer vision, particularly in applications requiring accurate and scalable representations across diverse perspectives. Current leading methods such as DUSt3R employ a…

Computer Vision and Pattern Recognition · Computer Science 2025-03-21 Jianing Yang , Alexander Sax , Kevin J. Liang , Mikael Henaff , Hao Tang , Ang Cao , Joyce Chai , Franziska Meier , Matt Feiszli

Understanding and reconstructing the complex geometry and motion of dynamic scenes from video remains a formidable challenge in computer vision. This paper introduces D4RT, a simple yet powerful feedforward model designed to efficiently…

Monocular dynamic video reconstruction faces significant challenges in dynamic human scenes due to geometric inconsistencies and resolution degradation issues. Existing methods lack 3D human structural understanding, producing geometrically…

Computer Vision and Pattern Recognition · Computer Science 2025-12-10 Weitao Xiong , Zhiyuan Yuan , Jiahao Lu , Chengfeng Zhao , Peng Li , Yuan Liu

Large Reconstruction Models (LRMs) have recently become a popular method for creating 3D foundational models. Training 3D reconstruction models with 2D visual data traditionally requires prior knowledge of camera poses for the training…

Computer Vision and Pattern Recognition · Computer Science 2025-03-11 Shiu-hong Kao , Xiao Li , Jinglu Wang , Yang Li , Chi-Keung Tang , Yu-Wing Tai , Yan Lu

In the realm of 3D reconstruction from 2D images, a persisting challenge is to achieve high-precision reconstructions devoid of 3D Ground Truth data reliance. We present UNeR3D, a pioneering unsupervised methodology that sets a new standard…

Computer Vision and Pattern Recognition · Computer Science 2023-12-13 Hongbin Lin , Juangui Xu , Qingfeng Xu , Zhengyu Hu , Handing Xu , Yunzhi Chen , Yongjun Hu , Zhenguo Nie

3D reconstruction of depth and motion from monocular video in dynamic environments is a highly ill-posed problem due to scale ambiguities when projecting to the 2D image domain. In this work, we investigate the performance of the current…

Computer Vision and Pattern Recognition · Computer Science 2022-01-24 Christian Homeyer , Oliver Lange , Christoph Schnörr

The recent paradigm shift in 3D vision led to the rise of foundation models with remarkable capabilities in 3D perception from uncalibrated images. However, extending these models to large-scale RGB stream 3D reconstruction remains…

Computer Vision and Pattern Recognition · Computer Science 2026-02-05 Leonid Antsfeld , Boris Chidlovskii , Yohann Cabon , Vincent Leroy , Jerome Revaud