English
Related papers

Related papers: MVRackLay: Monocular Multi-View Layout Estimation …

200 papers

Real-scale scene flow estimation has become increasingly important for 3D computer vision. Some works successfully estimate real-scale 3D scene flow with LiDAR. However, these ubiquitous and expensive sensors are still unlikely to be…

Computer Vision and Pattern Recognition · Computer Science 2021-11-25 Runfa Li , Truong Nguyen

We present a new paradigm for real-time object-oriented SLAM with a monocular camera. Contrary to previous approaches, that rely on object-level models, we construct category-level models from CAD collections which are now widely available.…

Robotics · Computer Science 2018-02-27 Parv Parkhiya , Rishabh Khawad , J. Krishna Murthy , Brojeshwar Bhowmick , K. Madhava Krishna

Traditional approaches for Visual Simultaneous Localization and Mapping (VSLAM) rely on low-level vision information for state estimation, such as handcrafted local features or the image gradient. While significant progress has been made…

Robotics · Computer Science 2021-08-05 Huaiyang Huang , Haoyang Ye , Yuxiang Sun , Lujia Wang , Ming Liu

This paper proposes a novel method to estimate the global scale of a 3D reconstructed model within a Kalman filtering-based monocular SLAM algorithm. Our Bayesian framework integrates height priors over the detected objects belonging to a…

Computer Vision and Pattern Recognition · Computer Science 2017-05-30 Edgar Sucar , Jean-Bernard Hayet

The rapid progress of Multimodal Large Language Models (MLLMs) has unlocked the potential for enhanced 3D scene understanding and spatial reasoning. A recent line of work explores learning spatial reasoning directly from multi-view images,…

Computer Vision and Pattern Recognition · Computer Science 2026-04-10 Kanghee Lee , Injae Lee , Minseok Kwak , Jungi Hong , Kwonyoung Ryu , Jaesik Park

This work proposes a new, online algorithm for estimating the local scale correction to apply to the output of a monocular SLAM system and obtain an as faithful as possible metric reconstruction of the 3D map and of the camera trajectory.…

Robotics · Computer Science 2017-11-09 Edgar Sucar , Jean-Bernard Hayet

This paper demonstrates a system capable of combining a sparse, indirect, monocular visual SLAM, with both offline and real-time Multi-View Stereo (MVS) reconstruction algorithms. This combination overcomes many obstacles encountered by…

Computer Vision and Pattern Recognition · Computer Science 2020-11-09 Fangwen Shu , Paul Lesur , Yaxu Xie , Alain Pagani , Didier Stricker

Recent unified 3D generation models have made remarkable progress in producing high-quality 3D assets from a single image. Notably, layout-aware approaches such as SAM3D can reconstruct multiple objects while preserving their spatial…

Computer Vision and Pattern Recognition · Computer Science 2026-04-10 Baicheng Li , Dong Wu , Jun Li , Shunkai Zhou , Zecui Zeng , Lusong Li , Hongbin Zha

We present Implicit-Scale 3D Reconstruction from Monocular Multi-Food Images, a benchmark dataset designed to advance geometry-based food portion estimation in realistic dining scenarios. Existing dietary assessment methods largely rely on…

Computer Vision and Pattern Recognition · Computer Science 2026-02-16 Yuhao Chen , Gautham Vinod , Siddeshwar Raghavan , Talha Ibn Mahmud , Bruce Coburn , Jinge Ma , Fengqing Zhu , Jiangpeng He

Pre-trained Vision-Language Models (VLMs), such as CLIP, have shown enhanced performance across a range of tasks that involve the integration of visual and linguistic modalities. When CLIP is used for depth estimation tasks, the patches,…

Computer Vision and Pattern Recognition · Computer Science 2023-11-03 Xueting Hu , Ce Zhang , Yi Zhang , Bowen Hai , Ke Yu , Zhihai He

Building structured 3D scene layouts from a single image requires reconciling visual observations with physical and spatial constraints, a challenge that is difficult to address with direct prediction alone. In this work, we formulate…

Computer Vision and Pattern Recognition · Computer Science 2026-05-26 Junwei Zhou , Yu-Wing Tai

Traditional monocular Visual Simultaneous Localization and Mapping (vSLAM) systems can be divided into three categories: those that use features, those that rely on the image itself, and hybrid models. In the case of feature-based methods,…

Robotics · Computer Science 2022-10-31 Andreas Georgis , Panagiotis Mermigkas , Petros Maragos

Reliable incremental estimation of camera poses and 3D reconstruction is key to enable various applications including robotics, interactive visualization, and augmented reality. However, this task is particularly challenging in dynamic…

Robotics · Computer Science 2025-12-09 Xingguang Zhong , Liren Jin , Marija Popović , Jens Behley , Cyrill Stachniss

With the popularity of monocular videos generated by video sharing and live broadcasting applications, reconstructing and editing dynamic scenes in stationary monocular cameras has become a special but anticipated technology. In contrast to…

Computer Vision and Pattern Recognition · Computer Science 2024-02-02 Weixing Xie , Xiao Dong , Yong Yang , Qiqin Lin , Jingze Chen , Junfeng Yao , Xiaohu Guo

Recent advances in structure-from-motion techniques are enabling many scientific fields to benefit from the routine creation of detailed 3D models. However, for a large number of applications, only a single camera is available, due to cost…

Computer Vision and Pattern Recognition · Computer Science 2019-06-20 Klemen Istenic , Nuno Gracias , Aurelien Arnaubec , Javier Escartin , Rafael Garcia

Although significant progress has been made in room layout estimation, most methods aim to reduce the loss in the 2D pixel coordinate rather than exploiting the room structure in the 3D space. Towards reconstructing the room layout in 3D,…

Computer Vision and Pattern Recognition · Computer Science 2021-04-06 Fu-En Wang , Yu-Hsuan Yeh , Min Sun , Wei-Chen Chiu , Yi-Hsuan Tsai

As the development of deep neural networks, 3D object recognition is becoming increasingly popular in computer vision community. Many multi-view based methods are proposed to improve the category recognition accuracy. These approaches…

Computer Vision and Pattern Recognition · Computer Science 2019-06-18 Qi Xuan , Fuxian Li , Yi Liu , Yun Xiang

The three-dimensional reconstruction of scenes from multiple views has made impressive strides in recent years, chiefly by methods correlating isolated feature points, intensities, or curvilinear structure. In the general setting, i.e.,…

Computer Vision and Pattern Recognition · Computer Science 2017-07-14 Anil Usumezbas , Ricardo Fabbri , Benjamin Kimia

This paper presents a hybrid real-time camera pose estimation framework with a novel partitioning scheme and introduces motion averaging to monocular Simultaneous Localization and Mapping (SLAM) systems. Breaking through the limitations of…

Computer Vision and Pattern Recognition · Computer Science 2020-11-04 Xinyi Li , Haibin Ling

A major element of depth perception and 3D understanding is the ability to predict the 3D layout of a scene and its contained objects for a novel pose. Indoor environments are particularly suitable for novel view prediction, since the set…

Computer Vision and Pattern Recognition · Computer Science 2018-08-13 Pulak Purkait , Ujwal Bonde , Christopher Zach