English
Related papers

Related papers: STATIC : Surface Temporal Affine for TIme Consiste…

200 papers

Mobile robots require accurate and robust depth measurements to understand and interact with the environment. While existing sensing modalities address this problem to some extent, recent research on monocular depth estimation has leveraged…

Robotics · Computer Science 2024-10-02 Marco Job , Thomas Stastny , Tim Kazik , Roland Siegwart , Michael Pantic

Self-supervised monocular depth estimation has been widely investigated to estimate depth images and relative poses from RGB images. This framework is attractive for researchers because the depth and pose networks can be trained from just…

Computer Vision and Pattern Recognition · Computer Science 2022-02-21 Noriaki Hirose , Kosuke Tahara

3D reconstruction of depth and motion from monocular video in dynamic environments is a highly ill-posed problem due to scale ambiguities when projecting to the 2D image domain. In this work, we investigate the performance of the current…

Computer Vision and Pattern Recognition · Computer Science 2022-01-24 Christian Homeyer , Oliver Lange , Christoph Schnörr

Consistent 3D geometry estimation from streaming RGB input is crucial for real-world applications such as autonomous driving, embodied AI, and large-scale reconstruction. While modern monocular geometry foundation models achieve strong…

Computer Vision and Pattern Recognition · Computer Science 2026-05-26 Xiaoyang Lyu , Muxin Liu , Xiaoshan Wu , Ruicheng Wang , Yi-Hua Huang , Yang-Tian Sun , Shaoshuai Shi , Xiaojuan Qi

Pre-trained on tremendous image-text pairs, vision-language models like CLIP have demonstrated promising zero-shot generalization across numerous image-based tasks. However, extending these capabilities to video tasks remains challenging…

Computer Vision and Pattern Recognition · Computer Science 2025-03-26 Zichen Liu , Kunlun Xu , Bing Su , Xu Zou , Yuxin Peng , Jiahuan Zhou

Accurate real depth annotations are difficult to acquire, needing the use of special devices such as a LiDAR sensor. Self-supervised methods try to overcome this problem by processing video or stereo sequences, which may not always be…

Computer Vision and Pattern Recognition · Computer Science 2020-09-04 Adrian Lopez-Rodriguez , Krystian Mikolajczyk

Recent advances in self-supervised learning havedemonstrated that it is possible to learn accurate monoculardepth reconstruction from raw video data, without using any 3Dground truth for supervision. However, in robotics…

Computer Vision and Pattern Recognition · Computer Science 2020-04-14 Robert McCraith , Lukas Neumann , Andrew Zisserman , Andrea Vedaldi

In self-supervised monocular depth estimation, the depth discontinuity and motion objects' artifacts are still challenging problems. Existing self-supervised methods usually utilize a single view to train the depth estimation network.…

Computer Vision and Pattern Recognition · Computer Science 2020-06-29 Jianrong Wang , Ge Zhang , Zhenyu Wu , XueWei Li , Li Liu

Monocular depth estimation using Convolutional Neural Networks (CNNs) has shown impressive performance in outdoor driving scenes. However, self-supervised learning of indoor depth from monocular sequences is quite challenging for…

Computer Vision and Pattern Recognition · Computer Science 2023-12-05 Chao Fan , Zhenyu Yin , Yue Li , Feiqing Zhang

Existing self-supervised monocular depth estimation methods can get rid of expensive annotations and achieve promising results. However, these methods suffer from severe performance degradation when directly adopting a model trained on a…

Computer Vision and Pattern Recognition · Computer Science 2022-07-27 Mu He , Le Hui , Yikai Bian , Jian Ren , Jin Xie , Jian Yang

Self-supervised learning for monocular depth estimation is widely investigated as an alternative to supervised learning approach, that requires a lot of ground truths. Previous works have successfully improved the accuracy of depth…

Computer Vision and Pattern Recognition · Computer Science 2020-11-25 Noriaki Hirose , Shun Taguchi , Keisuke Kawano , Satoshi Koide

Remarkable progress has been made in self-supervised monocular depth estimation (SS-MDE) by exploring cross-view consistency, e.g., photometric consistency and 3D point cloud consistency. However, they are very vulnerable to illumination…

Computer Vision and Pattern Recognition · Computer Science 2024-03-19 Haimei Zhao , Jing Zhang , Zhuo Chen , Bo Yuan , Dacheng Tao

In this paper, we tackle the problem of estimating the depth of a scene from a monocular video sequence. In particular, we handle challenging scenarios, such as non-translational camera motion and dynamic scenes, where traditional structure…

Computer Vision and Pattern Recognition · Computer Science 2015-11-20 Miaomiao Liu , Mathieu Salzmann , Xuming He

Accurate depth estimation from monocular videos remains challenging due to ambiguities inherent in single-view geometry, as crucial depth cues like stereopsis are absent. However, humans often perceive relative depth intuitively by…

Computer Vision and Pattern Recognition · Computer Science 2025-04-22 Seokju Cho , Jiahui Huang , Seungryong Kim , Joon-Young Lee

Spatial scene understanding, including monocular depth estimation, is an important problem in various applications, such as robotics and autonomous driving. While improvements in unsupervised monocular depth estimation have potentially…

Computer Vision and Pattern Recognition · Computer Science 2023-11-07 Hemang Chawla , Arnav Varma , Elahe Arani , Bahram Zonooz

Self-supervised depth estimation for indoor environments is more challenging than its outdoor counterpart in at least the following two aspects: (i) the depth range of indoor sequences varies a lot across different frames, making it…

Computer Vision and Pattern Recognition · Computer Science 2021-07-29 Pan Ji , Runze Li , Bir Bhanu , Yi Xu

Inverse Dynamics Models (IDMs) map visual observations to low-level action commands, serving as central components for data labeling and policy execution in embodied AI. However, their performance degrades severely under manipulator…

Robotics · Computer Science 2026-04-21 Kerui Li , Zhe Jing , Xiaofeng Wang , Zheng Zhu , Yukun Zhou , Guan Huang , Dongze Li , Qingkai Yang , Huaibo Huang

Existing audio-driven visual dubbing methods have achieved great success. Despite this, we observe that the semantic ambiguity between spatial and temporal domains significantly degrades the synthesis stability for the dynamic faces. We…

Computer Vision and Pattern Recognition · Computer Science 2025-04-01 Zijun Ding , Mingdie Xiong , Congcong Zhu , Jingrun Chen

Temporal human action detection aims to identify and localize action segments within untrimmed videos, serving as a pivotal task in video understanding. Despite the progress achieved by prior architectures like CNN and Transformer models,…

Computer Vision and Pattern Recognition · Computer Science 2026-04-13 Yicheng Qiu , Keiji Yanai

Self-supervised monocular depth estimation holds significant importance in the fields of autonomous driving and robotics. However, existing methods are typically trained and tested on standard datasets, overlooking the impact of various…

Computer Vision and Pattern Recognition · Computer Science 2024-03-19 Ziyang Song , Ruijie Zhu , Chuxin Wang , Jiacheng Deng , Jianfeng He , Tianzhu Zhang
‹ Prev 1 3 4 5 6 7 10 Next ›