English
Related papers

Related papers: MR.ScaleMaster: Scale-Consistent Collaborative Map…

200 papers

Cooperative Simultaneous Localization and Mapping (C-SLAM) enables multiple agents to work together in mapping unknown environments while simultaneously estimating their own positions. This approach enhances robustness, scalability, and…

Robotics · Computer Science 2025-08-28 Joshua Bird , Jan Blumenkamp , Amanda Prorok

Monocular depth estimation (MDE) has been widely adopted in the perception systems of autonomous vehicles and mobile robots. However, existing approaches often struggle to maintain temporal consistency in depth estimation across consecutive…

Computer Vision and Pattern Recognition · Computer Science 2026-04-03 Leezy Han , Seunggyu Kim , Dongseok Shim , Hyeonbeom Lee

Recent advancements in video diffusion models have shown exceptional abilities in simulating real-world dynamics and maintaining 3D consistency. This progress inspires us to investigate the potential of these models to ensure dynamic…

Computer Vision and Pattern Recognition · Computer Science 2024-12-11 Jianhong Bai , Menghan Xia , Xintao Wang , Ziyang Yuan , Xiao Fu , Zuozhu Liu , Haoji Hu , Pengfei Wan , Di Zhang

This paper presents a visual SLAM system that uses both points and lines for robust camera localization, and simultaneously performs a piece-wise planar reconstruction (PPR) of the environment to provide a structural map in real-time. One…

Computer Vision and Pattern Recognition · Computer Science 2023-01-31 Fangwen Shu , Jiaxuan Wang , Alain Pagani , Didier Stricker

Recent monocular 3D shape reconstruction methods have shown promising zero-shot results on object-segmented images without any occlusions. However, their effectiveness is significantly compromised in real-world conditions, due to imperfect…

Computer Vision and Pattern Recognition · Computer Science 2025-06-10 Junhyeong Cho , Kim Youwang , Hunmin Yang , Tae-Hyun Oh

In this paper, we propose a novel dense surfel mapping system that scales well in different environments with only CPU computation. Using a sparse SLAM system to estimate camera poses, the proposed mapping system can fuse intensity images…

Robotics · Computer Science 2019-09-11 Kaixuan Wang , Fei Gao , Shaojie Shen

The visual SLAM method is widely used for self-localization and mapping in complex environments. Visual-inertia SLAM, which combines a camera with IMU, can significantly improve the robustness and enable scale weak-visibility, whereas…

Robotics · Computer Science 2020-03-06 Peng Gang , Lu Zezao , Chen Bocheng , Chen Shanliang , He Dingxin

This paper shows that the autoregressive model is an effective and scalable monocular depth estimator. Our idea is simple: We tackle the monocular depth estimation (MDE) task with an autoregressive prediction paradigm, based on two core…

Computer Vision and Pattern Recognition · Computer Science 2025-03-20 Jinhong Wang , Jian Liu , Dongqi Tang , Weiqiang Wang , Wentong Li , Danny Chen , Jintai Chen , Jian Wu

Estimating depth from a single image is a challenging visual task. Compared to relative depth estimation, metric depth estimation attracts more attention due to its practical physical significance and critical applications in real-life…

Computer Vision and Pattern Recognition · Computer Science 2026-02-25 Ruijie Zhu , Chuxin Wang , Ziyang Song , Li Liu , Tianzhu Zhang , Yongdong Zhang

Depth map estimation from images is an important task in robotic systems. Existing methods can be categorized into two groups including multi-view stereo and monocular depth estimation. The former requires cameras to have large overlapping…

Computer Vision and Pattern Recognition · Computer Science 2022-10-06 Jialei Xu , Xianming Liu , Yuanchao Bai , Junjun Jiang , Kaixuan Wang , Xiaozhi Chen , Xiangyang Ji

Monocular 3D object detectors, while effective on data from one ego camera height, struggle with unseen or out-of-distribution camera heights. Existing methods often rely on Plucker embeddings, image transformations or data augmentation.…

Computer Vision and Pattern Recognition · Computer Science 2025-08-18 Abhinav Kumar , Yuliang Guo , Zhihao Zhang , Xinyu Huang , Liu Ren , Xiaoming Liu

Obtaining dense 3D reconstrution with low computational cost is one of the important goals in the field of SLAM. In this paper we propose a dense 3D reconstruction framework from monocular multispectral video sequences using jointly…

Computer Vision and Pattern Recognition · Computer Science 2018-07-09 Yuanhong Xu , Pei Dong , Junyu Dong , Lin Qi

Object SLAM introduces the concept of objects into Simultaneous Localization and Mapping (SLAM) and helps understand indoor scenes for mobile robots and object-level interactive applications. The state-of-art object SLAM systems face…

Robotics · Computer Science 2021-09-13 Ziwei Liao , Yutong Hu , Jiadong Zhang , Xianyu Qi , Xiaoyu Zhang , Wei Wang

One of the major challenges in Minimally Invasive Surgery (MIS) such as laparoscopy is the lack of depth perception. In recent years, laparoscopic scene tracking and surface reconstruction has been a focus of investigation to provide rich…

Computer Vision and Pattern Recognition · Computer Science 2017-03-06 Long Chen , Wen Tang , Nigel W. John , Tao Ruan Wan , Jian Jun Zhang

Capturing challenging human motions is critical for numerous applications, but it suffers from complex motion patterns and severe self-occlusion under the monocular setting. In this paper, we propose ChallenCap -- a template-based approach…

Computer Vision and Pattern Recognition · Computer Science 2021-03-30 Yannan He , Anqi Pang , Xin Chen , Han Liang , Minye Wu , Yuexin Ma , Lan Xu

Recovering multi-person 3D poses with absolute scales from a single RGB image is a challenging problem due to the inherent depth and scale ambiguity from a single view. Addressing this ambiguity requires to aggregate various cues over the…

Computer Vision and Pattern Recognition · Computer Science 2020-08-27 Jianan Zhen , Qi Fang , Jiaming Sun , Wentao Liu , Wei Jiang , Hujun Bao , Xiaowei Zhou

We introduce a novel framework for reconstructing dynamic human-object interactions from monocular video that overcomes challenges associated with occlusions and temporal inconsistencies. Traditional 3D reconstruction methods typically…

Computer Vision and Pattern Recognition · Computer Science 2025-09-16 Hyungjun Doh , Dong In Lee , Seunggeun Chi , Pin-Hao Huang , Kwonjoon Lee , Sangpil Kim , Karthik Ramani

We introduce MetricHMSR, a novel framework for recovering metric human meshes and 3D scenes from a single monocular image. Existing methods struggle to recover metric scale due to monocular scale ambiguity and weak-perspective camera…

Computer Vision and Pattern Recognition · Computer Science 2026-04-01 Chentao Song , He Zhang , Haolei Yuan , Haozhe Lin , Jianhua Tao , Hongwen Zhang , Tao Yu

Feedforward monocular face capture methods seek to reconstruct posed faces from a single image of a person. Current state of the art approaches have the ability to regress parametric 3D face models in real-time across a wide range of…

Computer Vision and Pattern Recognition · Computer Science 2024-09-13 Kelian Baert , Shrisha Bharadwaj , Fabien Castan , Benoit Maujean , Marc Christie , Victoria Abrevaya , Adnane Boukhayma

Recovering the metric 3D shape from a single image is particularly relevant for robotics and embodied intelligence applications, where accurate spatial understanding is crucial for navigation and interaction with environments. Usually, the…

Computer Vision and Pattern Recognition · Computer Science 2025-02-14 Chenghao Zhang , Lubin Fan , Shen Cao , Bojian Wu , Jieping Ye