English
Related papers

Related papers: Scale-aware Insertion of Virtual Objects in Monocu…

200 papers

Monocular depth estimation in the wild inherently predicts depth up to an unknown scale. To resolve scale ambiguity issue, we present a learning algorithm that leverages monocular simultaneous localization and mapping (SLAM) with…

Computer Vision and Pattern Recognition · Computer Science 2022-03-11 Jaehoon Choi , Dongki Jung , Yonghan Lee , Deokhwa Kim , Dinesh Manocha , Donghwan Lee

The rise of chronic diseases related to diet, such as obesity and diabetes, emphasizes the need for accurate monitoring of food intake. While AI-driven dietary assessment has made strides in recent years, the ill-posed nature of recovering…

Computer Vision and Pattern Recognition · Computer Science 2026-01-29 Gautham Vinod , Bruce Coburn , Siddeshwar Raghavan , Jiangpeng He , Fengqing Zhu

Monocular visual odometry (VO) has attracted extensive research attention by providing real-time vehicle motion from cost-effective camera images. However, state-of-the-art optimization-based monocular VO methods suffer from the scale…

Computer Vision and Pattern Recognition · Computer Science 2022-03-14 Sen Zhang , Jing Zhang , Dacheng Tao

This paper presents a new self-supervised system for learning to detect novel and previously unseen categories of objects in images. The proposed system receives as input several unlabeled videos of scenes containing various objects. The…

Computer Vision and Pattern Recognition · Computer Science 2021-08-25 Juntao Tan , Changkyu Song , Abdeslam Boularias

We introduce MetricHMSR, a novel framework for recovering metric human meshes and 3D scenes from a single monocular image. Existing methods struggle to recover metric scale due to monocular scale ambiguity and weak-perspective camera…

Computer Vision and Pattern Recognition · Computer Science 2026-04-01 Chentao Song , He Zhang , Haolei Yuan , Haozhe Lin , Jianhua Tao , Hongwen Zhang , Tao Yu

Self-supervised monocular depth estimation approaches either ignore independently moving objects in the scene or need a separate segmentation step to identify them. We propose MonoDepthSeg to jointly estimate depth and segment moving…

Computer Vision and Pattern Recognition · Computer Science 2021-10-22 Sadra Safadoust , Fatma Güney

In vision-enabled autonomous systems such as robots and autonomous cars, video object detection plays a crucial role, and both its speed and accuracy are important factors to provide reliable operation. The key insight we show in this paper…

Computer Vision and Pattern Recognition · Computer Science 2019-02-11 Ting-Wu Chin , Ruizhou Ding , Diana Marculescu

Autonomous robots often rely on monocular cameras for odometry estimation and navigation. However, the scale ambiguity problem presents a critical barrier to effective monocular visual odometry. In this paper, we present CodedVO, a novel…

Robotics · Computer Science 2024-07-26 Sachin Shah , Naitri Rajyaguru , Chahat Deep Singh , Christopher Metzler , Yiannis Aloimonos

Dense depth estimation is essential to scene-understanding for autonomous driving. However, recent self-supervised approaches on monocular videos suffer from scale-inconsistency across long sequences. Utilizing data from the ubiquitously…

Computer Vision and Pattern Recognition · Computer Science 2023-02-03 Hemang Chawla , Arnav Varma , Elahe Arani , Bahram Zonooz

Monocular visual SLAM enables 3D reconstruction from internet video and autonomous navigation on resource-constrained platforms, yet suffers from scale drift, i.e., the gradual divergence of estimated scale over long sequences. Existing…

Computer Vision and Pattern Recognition · Computer Science 2026-01-15 Yuchen Wu , Jiahe Li , Xiaohan Yu , Lina Yu , Jin Zheng , Xiao Bai

In the realm of computer vision, the perception and reconstruction of the 3D world through vision signals heavily rely on camera intrinsic parameters, which have long been a subject of intense research within the community. In practical…

Computer Vision and Pattern Recognition · Computer Science 2024-10-01 Xinyue Zhang , Jiaqi Yang , Xiangting Meng , Abdelrahman Mohamed , Laurent Kneip

Despite significant progress made in the past few years, challenges remain for depth estimation using a single monocular image. First, it is nontrivial to train a metric-depth prediction model that can generalize well to diverse scenes…

Computer Vision and Pattern Recognition · Computer Science 2022-09-07 Wei Yin , Jianming Zhang , Oliver Wang , Simon Niklaus , Simon Chen , Yifan Liu , Chunhua Shen

Realistic reconstruction of dynamic 4D scenes from monocular videos is essential for understanding the physical world. Despite recent progress in neural rendering, existing methods often struggle to recover accurate 3D geometry and…

Computer Vision and Pattern Recognition · Computer Science 2026-04-01 Haoran Zhou , Gim Hee Lee

Monocular visual localization plays a pivotal role in advanced driver assistance systems and autonomous driving by estimating a vehicle's ego-motion from a single pinhole camera. Nevertheless, conventional monocular visual odometry…

Computer Vision and Pattern Recognition · Computer Science 2025-03-07 Hui Zhang , Zhiyang Wu , Qianqian Shangguan , Kang An

Transparent object perception is indispensable for numerous robotic tasks. However, accurately segmenting and estimating the depth of transparent objects remain challenging due to complex optical properties. Existing methods primarily delve…

Computer Vision and Pattern Recognition · Computer Science 2025-03-04 Jiangyuan Liu , Hongxuan Ma , Yuxin Guo , Yuhao Zhao , Chi Zhang , Wei Sui , Wei Zou

The increasing accuracy reports of metric monocular depth estimation models lead to a growing interest from the automotive domain. Current model evaluations do not provide deeper insights into the models' performance, also in relation to…

Computer Vision and Pattern Recognition · Computer Science 2024-09-13 Tim Bader , Leon Eisemann , Adrian Pogorzelski , Namrata Jangid , Attila-Balazs Kis

Camera geo-localization from a monocular video is a fundamental task for video analysis and autonomous navigation. Although 3D reconstruction is a key technique to obtain camera poses, monocular 3D reconstruction in a large environment…

Computer Vision and Pattern Recognition · Computer Science 2018-08-28 Kazuya Iwami , Satoshi Ikehata , Kiyoharu Aizawa

The self-supervised loss formulation for jointly training depth and egomotion neural networks with monocular images is well studied and has demonstrated state-of-the-art accuracy. One of the main limitations of this approach, however, is…

Robotics · Computer Science 2022-05-03 Brandon Wagstaff , Jonathan Kelly

The objective of this paper is a model that is able to discover, track and segment multiple moving objects in a video. We make four contributions: First, we introduce an object-centric segmentation model with a depth-ordered layer…

Computer Vision and Pattern Recognition · Computer Science 2022-11-15 Junyu Xie , Weidi Xie , Andrew Zisserman

Monocular video human mesh recovery faces fundamental challenges in maintaining metric consistency and temporal stability due to inherent depth ambiguities and scale uncertainties. While existing methods rely primarily on RGB features and…

Computer Vision and Pattern Recognition · Computer Science 2026-02-05 Jiaxin Cen , Xudong Mao , Guanghui Yue , Wei Zhou , Ruomei Wang , Fan Zhou , Baoquan Zhao