English
Related papers

Related papers: MonoMAE: Enhancing Monocular 3D Detection through …

200 papers

Self-supervised monocular depth estimation, aiming to learn scene depths from single images in a self-supervised manner, has received much attention recently. In spite of recent efforts in this field, how to learn accurate scene depths and…

Computer Vision and Pattern Recognition · Computer Science 2022-09-16 Zhengming Zhou , Qiulei Dong

Accurate depth estimation is fundamental to 3D perception in autonomous driving, supporting tasks such as detection, tracking, and motion planning. However, monocular camera-based 3D detection suffers from depth ambiguity and reduced…

Computer Vision and Pattern Recognition · Computer Science 2026-01-28 Chen-Chou Lo , Patrick Vandewalle

Unsupervised learning of depth and ego-motion from unlabelled monocular videos has recently drawn great attention, which avoids the use of expensive ground truth in the supervised one. It achieves this by using the photometric errors…

Computer Vision and Pattern Recognition · Computer Science 2020-11-23 Hualie Jiang , Laiyan Ding , Zhenglong Sun , Rui Huang

"Masked Autoencoders (MAE) Are Scalable Vision Learners" revolutionizes the self-supervised learning method in that it not only achieves the state-of-the-art for image pre-training, but is also a milestone that bridges the gap between…

Computer Vision and Pattern Recognition · Computer Science 2022-02-10 Shuhao Cao , Peng Xu , David A. Clifton

The main challenge of monocular 3D object detection is the accurate localization of 3D center. Motivated by a new and strong observation that this challenge can be remedied by a 3D-space local-grid search scheme in an ideal case, we propose…

Computer Vision and Pattern Recognition · Computer Science 2023-04-05 Xianpeng Liu , Ce Zheng , Kelvin Cheng , Nan Xue , Guo-Jun Qi , Tianfu Wu

This paper investigates the geometric consistency for monocular 3D object detection, which suffers from the ill-posed depth estimation. We first conduct a thorough analysis to reveal how existing methods fail to consistently localize…

Computer Vision and Pattern Recognition · Computer Science 2022-05-24 Qing Lian , Botao Ye , Ruijia Xu , Weilong Yao , Tong Zhang

Spatial scene understanding, including monocular depth estimation, is an important problem in various applications, such as robotics and autonomous driving. While improvements in unsupervised monocular depth estimation have potentially…

Computer Vision and Pattern Recognition · Computer Science 2023-11-07 Hemang Chawla , Arnav Varma , Elahe Arani , Bahram Zonooz

Current perception models in autonomous driving heavily rely on large-scale labelled 3D data, which is both costly and time-consuming to annotate. This work proposes a solution to reduce the dependence on labelled 3D training data by…

Computer Vision and Pattern Recognition · Computer Science 2023-10-10 Chen Min , Xinli Xu , Dawei Zhao , Liang Xiao , Yiming Nie , Bin Dai

Monocular 3D object detection poses a significant challenge due to the lack of depth information in RGB images. Many existing methods strive to enhance the object depth estimation performance by allocating additional parameters for object…

Computer Vision and Pattern Recognition · Computer Science 2024-01-03 Wonhyeok Choi , Mingyu Shin , Sunghoon Im

Monocular image-based 3D perception has become an active research area in recent years owing to its applications in autonomous driving. Approaches to monocular 3D perception including detection and tracking, however, often yield inferior…

Computer Vision and Pattern Recognition · Computer Science 2022-06-09 Longlong Jing , Ruichi Yu , Henrik Kretzschmar , Kang Li , Charles R. Qi , Hang Zhao , Alper Ayvaci , Xu Chen , Dillon Cower , Yingwei Li , Yurong You , Han Deng , Congcong Li , Dragomir Anguelov

Estimating 3D orientation and translation of objects is essential for infrastructure-less autonomous navigation and driving. In case of monocular vision, successful methods have been mainly based on two ingredients: (i) a network generating…

Computer Vision and Pattern Recognition · Computer Science 2020-02-25 Zechen Liu , Zizhang Wu , Roland Tóth

Object pose estimation is a core perception task that enables, for example, object grasping and scene understanding. The widely available, inexpensive and high-resolution RGB sensors and CNNs that allow for fast inference based on this…

Monocular 3D lane detection is challenged by aleatoric uncertainty arising from inherent observation noise. Existing methods rely on simplified geometric assumptions, such as independent point predictions or global planar modeling, failing…

Computer Vision and Pattern Recognition · Computer Science 2025-11-18 Ruixin Liu , Zejian Yuan

We present a novel unsupervised learning framework for single view depth estimation using monocular videos. It is well known in 3D vision that enlarging the baseline can increase the depth estimation accuracy, and jointly optimizing a set…

Computer Vision and Pattern Recognition · Computer Science 2018-12-11 Lipu Zhou , Jiamin Ye , Montiel Abello , Shengze Wang , Michael Kaess

Today's state-of-the-art methods for 3D object detection are based on lidar, stereo, or monocular cameras. Lidar-based methods achieve the best accuracy, but have a large footprint, high cost, and mechanically-limited angular sampling…

Computer Vision and Pattern Recognition · Computer Science 2021-02-09 Frank Julca-Aguilar , Jason Taylor , Mario Bijelic , Fahim Mannan , Ethan Tseng , Felix Heide

Can a neural network estimate an object's dimension in the wild? In this paper, we propose a method and deep learning architecture to estimate the dimensions of a quadrilateral object of interest in videos using a monocular camera. The…

Computer Vision and Pattern Recognition · Computer Science 2022-12-29 Thariq Khalid , Mohammed Yahya Hakami , Riad Souissi

Previous works concerning single-view hand-held object reconstruction typically rely on supervision from 3D ground-truth models, which are hard to collect in real world. In contrast, readily accessible hand-object videos offer a promising…

Computer Vision and Pattern Recognition · Computer Science 2024-03-14 Chenyangguang Zhang , Guanlong Jiao , Yan Di , Gu Wang , Ziqin Huang , Ruida Zhang , Fabian Manhardt , Bowen Fu , Federico Tombari , Xiangyang Ji

Masked Autoencoder (MAE) is a notable method for self-supervised pretraining in visual representation learning. It operates by randomly masking image patches and reconstructing these masked patches using the unmasked ones. A key limitation…

Computer Vision and Pattern Recognition · Computer Science 2025-03-25 Han Guo , Ramtin Hosseini , Ruiyi Zhang , Sai Ashish Somayajula , Ranak Roy Chowdhury , Rajesh K. Gupta , Pengtao Xie

We present a novel monocular localization framework by jointly training deep learning-based depth prediction and Bayesian filtering-based pose reasoning. The proposed cross-modal framework significantly outperforms deep learning-only…

Computer Vision and Pattern Recognition · Computer Science 2022-10-28 Priyesh Shukla , Sureshkumar S. , Alex C. Stutts , Sathya Ravi , Theja Tulabandhula , Amit R. Trivedi

We present a method for jointly training the estimation of depth, ego-motion, and a dense 3D translation field of objects relative to the scene, with monocular photometric consistency being the sole source of supervision. We show that this…

Computer Vision and Pattern Recognition · Computer Science 2020-11-10 Hanhan Li , Ariel Gordon , Hang Zhao , Vincent Casser , Anelia Angelova