English
Related papers

Related papers: MVRackLay: Monocular Multi-View Layout Estimation …

200 papers

Given a monocular colour image of a warehouse rack, we aim to predict the bird's-eye view layout for each shelf in the rack, which we term as multi-layer layout prediction. To this end, we present RackLay, a deep neural network for…

Computer Vision and Pattern Recognition · Computer Science 2021-10-29 Meher Shashwat Nigam , Avinash Prabhu , Anurag Sahu , Puru Gupta , Tanvi Karandikar , N. Sai Shankar , Ravi Kiran Sarvadevabhatla , K. Madhava Krishna

We present 360-DFPE, a sequential floor plan estimation method that directly takes 360-images as input without relying on active sensors or 3D information. Our approach leverages a loosely coupled integration between a monocular visual SLAM…

Computer Vision and Pattern Recognition · Computer Science 2022-05-10 Bolivar Solarte , Yueh-Cheng Liu , Chin-Hsuan Wu , Yi-Hsuan Tsai , Min Sun

Spherical cameras capture scenes in a holistic manner and have been used for room layout estimation. Recently, with the availability of appropriate datasets, there has also been progress in depth estimation from a single omnidirectional…

Computer Vision and Pattern Recognition · Computer Science 2022-06-24 Nikolaos Zioulis , Federico Alvarez , Dimitrios Zarpalas , Petros Daras

We present MVLayoutNet, an end-to-end network for holistic 3D reconstruction from multi-view panoramas. Our core contribution is to seamlessly combine learned monocular layout estimation and multi-view stereo (MVS) for accurate layout…

Computer Vision and Pattern Recognition · Computer Science 2021-12-14 Zhihua Hu , Bo Duan , Yanfeng Zhang , Mingwei Sun , Jingwei Huang

In this paper, we present a monocular Simultaneous Localization and Mapping (SLAM) algorithm using high-level object and plane landmarks. The built map is denser, more compact and semantic meaningful compared to feature point based SLAM. We…

Robotics · Computer Science 2019-07-01 Shichao Yang , Sebastian Scherer

Multi-camera systems have been shown to improve the accuracy and robustness of SLAM estimates, yet state-of-the-art SLAM systems predominantly support monocular or stereo setups. This paper presents a generic sparse visual SLAM framework…

Dense and accurate 3D mapping from a monocular sequence is a key technology for several applications and still an open research area. This paper leverages recent results on single-view CNN-based depth estimation and fuses them with…

Computer Vision and Pattern Recognition · Computer Science 2017-06-28 José M. Fácil , Alejo Concha , Luis Montesano , Javier Civera

Pre-trained general-purpose Vision-Language Models (VLM) hold the potential to enhance intuitive human-machine interactions due to their rich world knowledge and 2D object detection capabilities. However, VLMs for 3D coordinates detection…

Computer Vision and Pattern Recognition · Computer Science 2026-03-03 Ari Wahl , Dorian Gawlinski , David Przewozny , Paul Chojecki , Felix Bießmann , Sebastian Bosse

In this paper, we address the novel, highly challenging problem of estimating the layout of a complex urban driving scenario. Given a single color image captured from a driving platform, we aim to predict the bird's-eye view layout of the…

Computer Vision and Pattern Recognition · Computer Science 2020-02-21 Kaustubh Mani , Swapnil Daga , Shubhika Garg , N. Sai Shankar , Krishna Murthy Jatavallabhula , K. Madhava Krishna

Monocular 3D object detection is a fundamental yet challenging task in 3D scene understanding. Existing approaches heavily depend on supervised learning with extensive 3D annotations, which are often acquired from LiDAR point clouds through…

Computer Vision and Pattern Recognition · Computer Science 2025-12-02 Zihua Liu , Hiroki Sakuma , Masatoshi Okutomi

We propose a scene-level inverse rendering framework that uses multi-view images to decompose the scene into geometry, a SVBRDF, and 3D spatially-varying lighting. Because multi-view images provide a variety of information about the scene,…

Computer Vision and Pattern Recognition · Computer Science 2023-03-28 JunYong Choi , SeokYeong Lee , Haesol Park , Seung-Won Jung , Ig-Jae Kim , Junghyun Cho

Existing inverse physics methods recover physical parameters from multi-view videos, where geometric constraints across views resolve scale and 3D structure. In monocular settings, however, such constraints are absent, leading to severe…

Computer Vision and Pattern Recognition · Computer Science 2026-05-29 Daniel Rho , Jun Myeong Choi , Matthew Thornton , Biswadip Dey , Roni Sengupta

Visual SLAM systems targeting static scenes have been developed with satisfactory accuracy and robustness. Dynamic 3D object tracking has then become a significant capability in visual SLAM with the requirement of understanding dynamic…

Computer Vision and Pattern Recognition · Computer Science 2022-10-06 Hanwei Zhang , Hideaki Uchiyama , Shintaro Ono , Hiroshi Kawasaki

Given an image or a video captured from a monocular camera, amodal layout estimation is the task of predicting semantics and occupancy in bird's eye view. The term amodal implies we also reason about entities in the scene that are occluded…

Robotics · Computer Science 2021-08-23 Kaustubh Mani , N. Sai Shankar , Krishna Murthy Jatavallabhula , K. Madhava Krishna

Recent work has shown impressive localization performance using only images of ground textures taken with a downward facing monocular camera. This provides a reliable navigation method that is robust to feature sparse environments and…

Robotics · Computer Science 2023-03-13 Kyle M. Hart , Brendan Englot , Ryan P. O'Shea , John D. Kelly , David Martinez

Laparoscopic video tracking primarily focuses on two target types: surgical instruments and anatomy. The former could be used for skill assessment, while the latter is necessary for the projection of virtual overlays. Where instrument and…

Computer Vision and Pattern Recognition · Computer Science 2024-03-29 Beerend G. A. Gerats , Jelmer M. Wolterink , Seb P. Mol , Ivo A. M. J. Broeders

Recently there has been a growing interest in category-level object pose and size estimation, and prevailing methods commonly rely on single view RGB-D images. However, one disadvantage of such methods is that they require accurate depth…

Computer Vision and Pattern Recognition · Computer Science 2024-03-25 Jiaqi Yang , Yucong Chen , Xiangting Meng , Chenxin Yan , Min Li , Ran Cheng , Lige Liu , Tao Sun , Laurent Kneip

We present a novel method to reconstruct the 3D layout of a room (walls, floors, ceilings) from a single perspective view in challenging conditions, by contrast with previous single-view methods restricted to cuboid-shaped layouts. This…

Computer Vision and Pattern Recognition · Computer Science 2020-07-22 Sinisa Stekovic , Shreyas Hampali , Mahdi Rad , Sayan Deb Sarkar , Friedrich Fraundorfer , Vincent Lepetit

Fiducial markers can encode rich information about the environment and can aid Visual SLAM (VSLAM) approaches in reconstructing maps with practical semantic information. Current marker-based VSLAM approaches mainly utilize markers for…

Robotics · Computer Science 2023-12-27 Ali Tourani , Hriday Bavle , Jose Luis Sanchez-Lopez , Rafael Munoz Salinas , Holger Voos

Spatial reasoning from monocular images is essential for autonomous driving, yet current Vision-Language Models (VLMs) still struggle with fine-grained geometric perception, particularly under large scale variation and ambiguous object…

Computer Vision and Pattern Recognition · Computer Science 2026-03-10 Yanchun Cheng , Rundong Wang , Xulei Yang , Alok Prakash , Daniela Rus , Marcelo H Ang , ShiJie Li
‹ Prev 1 2 3 10 Next ›