English
Related papers

Related papers: MVRackLay: Monocular Multi-View Layout Estimation …

200 papers

Object SLAM uses additional semantic information to detect and map objects in the scene, in order to improve the system's perception and map representation capabilities. Quadrics and cubes are often used to represent objects, but their…

Robotics · Computer Science 2022-09-23 Xiao Han , Lu Yang

In recent years, Multi-View Clustering (MVC) has attracted increasing attention for its potential to reduce the annotation burden associated with large datasets. The aim of MVC is to exploit the inherent consistency and complementarity…

Machine Learning · Computer Science 2024-07-12 Zhangci Xiong , Meng Cao

Vision-language models (VLMs) have achieved impressive results on single-view vision tasks, but lack the multi-view spatial reasoning capabilities essential for embodied AI systems to understand 3D environments and manipulate objects across…

Computer Vision and Pattern Recognition · Computer Science 2026-03-31 Suchae Jeong , Jaehwi Song , Haeone Lee , Hanna Kim , Jian Kim , Dongjun Lee , Dong Kyu Shin , Changyeon Kim , Dongyoon Hahm , Woogyeol Jin , Juheon Choi , Kimin Lee

3D Visual Grounding (3DVG) aims to localize objects in 3D scenes using natural language descriptions. Although supervised methods achieve higher accuracy in constrained settings, zero-shot 3DVG holds greater promise for real-world…

Computer Vision and Pattern Recognition · Computer Science 2025-08-29 Jiawen Lin , Shiran Bian , Yihang Zhu , Wenbin Tan , Yachao Zhang , Yuan Xie , Yanyun Qu

The goal of this paper is to compare surface-based and volumetric 3D object shape representations, as well as viewer-centered and object-centered reference frames for single-view 3D shape prediction. We propose a new algorithm for…

Computer Vision and Pattern Recognition · Computer Science 2018-06-13 Daeyun Shin , Charless C. Fowlkes , Derek Hoiem

This paper aims to design a 3D object detection model from 2D images taken by monocular cameras by combining the estimated bird's-eye view elevation map and the deep representation of object features. The proposed model has a pre-trained…

Computer Vision and Pattern Recognition · Computer Science 2020-11-25 Ali Babolhavaeji , Mohammad Fanaei

Vision--language models reliably name objects in a scene, but do they represent the 3D layout those objects inhabit? We introduce a 3,034-sample human-curated benchmark targeting three components of spatial understanding: depth-ordered…

Computer Vision and Pattern Recognition · Computer Science 2026-05-21 Animesh Maheshwari , Divyansh Sahu , Nishit Verma

Monocular simultaneous localization and mapping (SLAM) algorithms estimate drone poses and build a 3D map using a single camera. Current algorithms include sparse methods that lack detailed geometry, while learning-driven approaches produce…

Robotics · Computer Science 2025-11-25 Jeryes Danial , Yosi Ben Asher , Itzik Klein

Current 3D layout estimation models are primarily trained on synthetic datasets containing simple single room or single floor environments. As a consequence, they cannot natively handle large multi floor buildings and require scenes to be…

Computer Vision and Pattern Recognition · Computer Science 2025-12-03 Valentin Bieri , Marie-Julie Rakotosaona , Keisuke Tateno , Francis Engelmann , Leonidas Guibas

Monocular 3D detection is a challenging task due to the lack of accurate 3D information. Existing approaches typically rely on geometry constraints and dense depth estimates to facilitate the learning, but often fail to fully exploit the…

Computer Vision and Pattern Recognition · Computer Science 2023-05-26 Liang Peng , Junkai Xu , Haoran Cheng , Zheng Yang , Xiaopei Wu , Wei Qian , Wenxiao Wang , Boxi Wu , Deng Cai

Neural networks have shown great success in extracting geometric information from color images. Especially, monocular depth estimation networks are increasingly reliable in real-world scenes. In this work we investigate the applicability of…

Computer Vision and Pattern Recognition · Computer Science 2023-02-21 Dominik Engel , Sebastian Hartwig , Timo Ropinski

Vision-Language Navigation (VLN) tasks often leverage panoramic RGB and depth inputs to provide rich spatial cues for action planning, but these sensors can be costly or less accessible in real-world deployments. Recent approaches based on…

Computer Vision and Pattern Recognition · Computer Science 2025-12-01 Shuo Wang , Yongcai Wang , Zhaoxin Fan , Yucheng Wang , Maiyue Chen , Kaihui Wang , Zhizhong Su , Wanting Li , Xudong Cai , Yeying Jin , Deying Li

Perception is a fundamental task in the field of computer vision, encompassing a diverse set of subtasks that can be systematically categorized into four distinct groups based on two dimensions: prediction type and instruction type.…

Computer Vision and Pattern Recognition · Computer Science 2025-07-23 Wentao Xiang , Haoxian Tan , Cong Wei , Yujie Zhong , Dengjie Li , Yujiu Yang

Vision-language models (VLMs) achieve strong performance on spatial reasoning benchmarks, yet it remains unclear whether this reflects structured 3D understanding or reliance on statistical shortcuts in natural images. We introduce a…

Computer Vision and Pattern Recognition · Computer Science 2026-05-29 Cheolhong Min , Jaeyun Jung , Daeun Lee , Hyeonseong Jeon , Yu Su , Jonathan Tremblay , Chan Hee Song , Jaesik Park

We present ViSTA-SLAM as a real-time monocular visual SLAM system that operates without requiring camera intrinsics, making it broadly applicable across diverse camera setups. At its core, the system employs a lightweight symmetric two-view…

Computer Vision and Pattern Recognition · Computer Science 2026-01-07 Ganlin Zhang , Shenhan Qian , Xi Wang , Daniel Cremers

Accurate 3D lane estimation is crucial for ensuring safety in autonomous driving. However, prevailing monocular techniques suffer from depth loss and lighting variations, hampering accurate 3D lane detection. In contrast, LiDAR points offer…

Computer Vision and Pattern Recognition · Computer Science 2024-06-25 Yueru Luo , Shuguang Cui , Zhen Li

Deep convolutional neural networks achieve remarkable visual recognition performance, at the cost of high computational complexity. In this paper, we have a new design of efficient convolutional layers based on three schemes. The 3D…

Computer Vision and Pattern Recognition · Computer Science 2017-01-25 Min Wang , Baoyuan Liu , Hassan Foroosh

Detecting and localizing objects in the real 3D space, which plays a crucial role in scene understanding, is particularly challenging given only a monocular image due to the geometric information loss during imagery projection. We propose…

Computer Vision and Pattern Recognition · Computer Science 2021-04-20 Zengyi Qin , Jinglu Wang , Yan Lu

A monocular 3D object tracking system generally has only up-to-scale pose estimation results without any prior knowledge of the tracked object. In this paper, we propose a novel idea to recover the metric scale of an arbitrary dynamic…

Robotics · Computer Science 2018-08-22 Kejie Qiu , Tong Qin , Hongwen Xie , Shaojie Shen

We introduce MulayCap, a novel human performance capture method using a monocular video camera without the need for pre-scanning. The method uses "multi-layer" representations for geometry reconstruction and texture rendering, respectively.…

Computer Vision and Pattern Recognition · Computer Science 2020-10-05 Zhaoqi Su , Weilin Wan , Tao Yu , Lingjie Liu , Lu Fang , Wenping Wang , Yebin Liu