English
Related papers

Related papers: Learning to Fuse Monocular and Multi-view Cues for…

200 papers

Object pose estimation is a fundamental task in 3D vision with applications in robotics, AR/VR, and scene understanding. We address the challenge of category-level 9-DoF pose estimation (6D pose + 3Dsize) from RGB-D input, without relying…

Computer Vision and Pattern Recognition · Computer Science 2026-03-31 Rachit Agarwal , Abhishek Joshi , Sathish Chalasani , Woo Jin Kim

Self-supervised monocular depth estimation has been widely studied, owing to its practical importance and recent promising improvements. However, most works suffer from limited supervision of photometric consistency, especially in weak…

Computer Vision and Pattern Recognition · Computer Science 2021-08-20 Hyunyoung Jung , Eunhyeok Park , Sungjoo Yoo

Robust three-dimensional scene understanding is now an ever-growing area of research highly relevant in many real-world applications such as autonomous driving and robotic navigation. In this paper, we propose a multi-task learning-based…

Computer Vision and Pattern Recognition · Computer Science 2019-08-16 Amir Atapour-Abarghouei , Toby P. Breckon

Estimating 3D scene flow from a sequence of monocular images has been gaining increased attention due to the simple, economical capture setup. Owing to the severe ill-posedness of the problem, the accuracy of current methods has been…

Computer Vision and Pattern Recognition · Computer Science 2021-05-06 Junhwa Hur , Stefan Roth

In this letter, we propose a new method, Multi-Clue Gaze (MCGaze), to facilitate video gaze estimation via capturing spatial-temporal interaction context among head, face, and eye in an end-to-end learning way, which has not been well…

Computer Vision and Pattern Recognition · Computer Science 2024-01-01 Yiran Guan , Zhuoguang Chen , Wenzheng Zeng , Zhiguo Cao , Yang Xiao

Depth map estimation from images is an important task in robotic systems. Existing methods can be categorized into two groups including multi-view stereo and monocular depth estimation. The former requires cameras to have large overlapping…

Computer Vision and Pattern Recognition · Computer Science 2022-10-06 Jialei Xu , Xianming Liu , Yuanchao Bai , Junjun Jiang , Kaixuan Wang , Xiaozhi Chen , Xiangyang Ji

A well-known challenge in applying deep-learning methods to omnidirectional images is spherical distortion. In dense regression tasks such as depth estimation, where structural details are required, using a vanilla CNN layer on the…

Computer Vision and Pattern Recognition · Computer Science 2022-03-30 Yuyan Li , Yuliang Guo , Zhixin Yan , Xinyu Huang , Ye Duan , Liu Ren

Depth estimation from images serves as the fundamental step of 3D perception for autonomous driving and is an economical alternative to expensive depth sensors like LiDAR. The temporal photometric constraints enables self-supervised depth…

Computer Vision and Pattern Recognition · Computer Science 2022-09-21 Yi Wei , Linqing Zhao , Wenzhao Zheng , Zheng Zhu , Yongming Rao , Guan Huang , Jiwen Lu , Jie Zhou

This work delves into unsupervised monocular depth estimation in endoscopy, which leverages adjacent frames to establish a supervisory signal during the training phase. For many clinical applications, e.g., surgical navigation, temporally…

Computer Vision and Pattern Recognition · Computer Science 2025-02-18 Shuwei Shao , Zhongcai Pei , Weihai Chen , Xingming Wu , Zhong Liu

Accurate depth estimation is at the core of many applications in computer graphics, vision, and robotics. Current state-of-the-art monocular depth estimators, trained on extensive datasets, generalize well but lack 3D consistency needed for…

Computer Vision and Pattern Recognition · Computer Science 2025-11-26 Laura Fink , Linus Franke , Bernhard Egger , Joachim Keinert , Marc Stamminger

Scene flow estimation is an extremely important task in computer vision to support the perception of dynamic changes in the scene. For robust scene flow, learning-based approaches have recently achieved impressive results using either…

Computer Vision and Pattern Recognition · Computer Science 2026-02-26 Rajai Alhimdiat , Ramy Battrawy , René Schuster , Didier Stricker , Wesam Ashour

A monocular 3D object tracking system generally has only up-to-scale pose estimation results without any prior knowledge of the tracked object. In this paper, we propose a novel idea to recover the metric scale of an arbitrary dynamic…

Robotics · Computer Science 2018-08-22 Kejie Qiu , Tong Qin , Hongwen Xie , Shaojie Shen

Effective feature fusion of multispectral images plays a crucial role in multi-spectral object detection. Previous studies have demonstrated the effectiveness of feature fusion using convolutional neural networks, but these methods are…

Computer Vision and Pattern Recognition · Computer Science 2023-08-16 Jifeng Shen , Yifei Chen , Yue Liu , Xin Zuo , Heng Fan , Wankou Yang

In recent years, neural implicit surface reconstruction methods have become popular for multi-view 3D reconstruction. In contrast to traditional multi-view stereo methods, these approaches tend to produce smoother and more complete…

Computer Vision and Pattern Recognition · Computer Science 2022-10-13 Zehao Yu , Songyou Peng , Michael Niemeyer , Torsten Sattler , Andreas Geiger

In this paper, we present a new method for multi-view geometric reconstruction. In recent years, large vision models have rapidly developed, performing excellently across various tasks and demonstrating remarkable generalization…

Computer Vision and Pattern Recognition · Computer Science 2025-03-19 Haoyu Guo , He Zhu , Sida Peng , Haotong Lin , Yunzhi Yan , Tao Xie , Wenguan Wang , Xiaowei Zhou , Hujun Bao

Monocular image-based 3D perception has become an active research area in recent years owing to its applications in autonomous driving. Approaches to monocular 3D perception including detection and tracking, however, often yield inferior…

Computer Vision and Pattern Recognition · Computer Science 2022-06-09 Longlong Jing , Ruichi Yu , Henrik Kretzschmar , Kang Li , Charles R. Qi , Hang Zhao , Alper Ayvaci , Xu Chen , Dillon Cower , Yingwei Li , Yurong You , Han Deng , Congcong Li , Dragomir Anguelov

Multi-modal semantic understanding requires integrating information from different modalities to extract users' real intention behind words. Most previous work applies a dual-encoder structure to separately encode image and text, but fails…

Computation and Language · Computer Science 2024-03-12 Ming Zhang , Ke Chang , Yunfang Wu

Human visual system relies on both binocular stereo cues and monocular focusness cues to gain effective 3D perception. In computer vision, the two problems are traditionally solved in separate tracks. In this paper, we present a unified…

Computer Vision and Pattern Recognition · Computer Science 2020-08-11 Xinqing Guo , Zhang Chen , Siyuan Li , Yang Yang , Jingyi Yu

Although deep neural networks have been widely applied to computer vision problems, extending them into multiview depth estimation is non-trivial. In this paper, we present MVDepthNet, a convolutional network to solve the depth estimation…

Robotics · Computer Science 2018-07-24 Kaixuan Wang , Shaojie Shen

Incrementally recovering 3D dense structures from monocular videos is of paramount importance since it enables various robotics and AR applications. Feature volumes have recently been shown to enable efficient and accurate incremental dense…

Computer Vision and Pattern Recognition · Computer Science 2023-05-25 Xingxing Zuo , Nan Yang , Nathaniel Merrill , Binbin Xu , Stefan Leutenegger