English
Related papers

Related papers: Mono-Hydra++: Real-Time Monocular Scene Graph Cons…

200 papers

Accurate 3D object detection is crucial to autonomous driving. Though LiDAR-based detectors have achieved impressive performance, the high cost of LiDAR sensors precludes their widespread adoption in affordable vehicles. Camera-based…

Computer Vision and Pattern Recognition · Computer Science 2024-04-11 Yurong You , Cheng Perng Phoo , Carlos Andres Diaz-Ruiz , Katie Z Luo , Wei-Lun Chao , Mark Campbell , Bharath Hariharan , Kilian Q Weinberger

The task of 3D semantic scene completion using monocular cameras is gaining significant attention in the field of autonomous driving. This task aims to predict the occupancy status and semantic labels of each voxel in a 3D scene from…

Computer Vision and Pattern Recognition · Computer Science 2024-11-27 Jiawei Yao , Jusheng Zhang , Xiaochao Pan , Tong Wu , Canran Xiao

Multi-frame methods improve monocular depth estimation over single-frame approaches by aggregating spatial-temporal information via feature matching. However, the spatial-temporal feature leads to accuracy degradation in dynamic scenes. To…

Computer Vision and Pattern Recognition · Computer Science 2023-12-20 Jiquan Zhong , Xiaolin Huang , Xiao Yu

Estimating the 3D position and orientation of objects in the environment with a single RGB camera is a critical and challenging task for low-cost urban autonomous driving and mobile robots. Most of the existing algorithms are based on the…

Computer Vision and Pattern Recognition · Computer Science 2021-02-02 Yuxuan Liu , Yuan Yixuan , Ming Liu

Monocular 3D object detection is an important yet challenging task in autonomous driving. Some existing methods leverage depth information from an off-the-shelf depth estimator to assist 3D detection, but suffer from the additional…

Computer Vision and Pattern Recognition · Computer Science 2022-03-29 Kuan-Chih Huang , Tsung-Han Wu , Hung-Ting Su , Winston H. Hsu

Detecting and localizing glass in 3D environments poses significant challenges for visual perception systems, as the optical properties of glass often hinder conventional sensors from accurately distinguishing glass surfaces. The lack of…

Robotics · Computer Science 2025-09-09 Kai Zhang , Guoyang Zhao , Jianxing Shi , Bonan Liu , Weiqing Qi , Jun Ma

Large Language Models (LLMs) can help robots reason about abstract task specifications. This requires augmenting classical representations of the environment used by robots, such as point-clouds and meshes, with natural language-based…

Robotics · Computer Science 2026-03-11 Christopher D. Hsu , Pratik Chaudhari

Monocular 3D object detection (Mono3D) in mobile settings (e.g., on a vehicle, a drone, or a robot) is an important yet challenging task. Due to the near-far disparity phenomenon of monocular vision and the ever-changing camera pose, it is…

Computer Vision and Pattern Recognition · Computer Science 2023-03-27 Yunsong Zhou , Quan Liu , Hongzi Zhu , Yunzhe Li , Shan Chang , Minyi Guo

Recent work has shown impressive localization performance using only images of ground textures taken with a downward facing monocular camera. This provides a reliable navigation method that is robust to feature sparse environments and…

Robotics · Computer Science 2023-03-13 Kyle M. Hart , Brendan Englot , Ryan P. O'Shea , John D. Kelly , David Martinez

Pseudo-LiDAR 3D detectors have made remarkable progress in monocular 3D detection by enhancing the capability of perceiving depth with depth estimation networks, and using LiDAR-based 3D detection architectures. The advanced stereo 3D…

Computer Vision and Pattern Recognition · Computer Science 2022-03-07 Yi-Nan Chen , Hang Dai , Yong Ding

Efficient target localization and autonomous navigation in complex environments are fundamental to real-world embodied applications. While recent advances in multimodal foundation models have enabled zero-shot object goal navigation,…

Robotics · Computer Science 2026-04-02 Ming-Ming Yu , Yi Chen , Börje F. Karlsson , Wenjun Wu

Simultaneous Localization and Mapping (SLAM) systems are fundamental building blocks for any autonomous robot navigating in unknown environments. The SLAM implementation heavily depends on the sensor modality employed on the mobile…

Learning to predict scene depth from RGB inputs is a challenging task both for indoor and outdoor robot navigation. In this work we address unsupervised learning of scene depth and robot ego-motion where supervision is provided by monocular…

Computer Vision and Pattern Recognition · Computer Science 2018-11-16 Vincent Casser , Soeren Pirk , Reza Mahjourian , Anelia Angelova

Mobile robots require accurate and robust depth measurements to understand and interact with the environment. While existing sensing modalities address this problem to some extent, recent research on monocular depth estimation has leveraged…

Robotics · Computer Science 2024-10-02 Marco Job , Thomas Stastny , Tim Kazik , Roland Siegwart , Michael Pantic

Scene understanding is paramount in robotics, self-navigation, augmented reality, and many other fields. To fully accomplish this task, an autonomous agent has to infer the 3D structure of the sensed scene (to know where it looks at) and…

Computer Vision and Pattern Recognition · Computer Science 2020-02-26 Pier Luigi Dovesi , Matteo Poggi , Lorenzo Andraghetti , Miquel Martí , Hedvig Kjellström , Alessandro Pieropan , Stefano Mattoccia

Reconstructing accurate 3D scenes from images is a long-standing vision task. Due to the ill-posedness of the single-image reconstruction problem, most well-established methods are built upon multi-view geometry. State-of-the-art (SOTA)…

Computer Vision and Pattern Recognition · Computer Science 2023-07-21 Wei Yin , Chi Zhang , Hao Chen , Zhipeng Cai , Gang Yu , Kaixuan Wang , Xiaozhi Chen , Chunhua Shen

Steering estimation is a critical task in autonomous driving, traditionally relying on 2D image-based models. In this work, we explore the advantages of incorporating 3D spatial information through hybrid architectures that combine 3D…

Computer Vision and Pattern Recognition · Computer Science 2025-03-24 Fouad Makiyeh , Huy-Dung Nguyen , Patrick Chareyre , Ramin Hasani , Marc Blanchon , Daniela Rus

This paper presents a compact and accurate representation of 3D scenes that are observed by a LiDAR sensor and a monocular camera. The proposed method is based on the well-established Stixel model originally developed for stereo vision…

Computer Vision and Pattern Recognition · Computer Science 2018-09-28 Florian Piewak , Peter Pinggera , Markus Enzweiler , David Pfeiffer , Marius Zöllner

Existing simultaneous localization and mapping (SLAM) algorithms are not robust in challenging low-texture environments because there are only few salient features. The resulting sparse or semi-dense map also conveys little information for…

Computer Vision and Pattern Recognition · Computer Science 2017-03-22 Shichao Yang , Yu Song , Michael Kaess , Sebastian Scherer

Monocular 3D object detection (Mono3D) aims to infer object locations and dimensions in 3D space from a single RGB image. Despite recent progress, existing methods remain highly sensitive to camera intrinsics and struggle to generalize…

Computer Vision and Pattern Recognition · Computer Science 2026-03-31 Zhihao Zhang , Abhinav Kumar , Xiaoming Liu