English
Related papers

Related papers: Empowering Dynamic Urban Navigation with Stereo an…

200 papers

Monocular depth estimation has been actively studied in fields such as robot vision, autonomous driving, and 3D scene understanding. Given a sequence of color images, unsupervised learning methods based on the framework of…

Computer Vision and Pattern Recognition · Computer Science 2022-12-06 Songlin Wei , Guodong Chen , Wenzheng Chi , Zhenhua Wang , Lining Sun

Vision-Language Models (VLMs) have advanced rapidly in multimodal perception and language understanding, yet it remains unclear whether they can reliably ground language into spatially coherent, plausibly executable actions in 3D digital…

Computer Vision and Pattern Recognition · Computer Science 2026-05-12 Niyati Rawal , Sushant Ravva , Shah Alam Abir , Saksham Jain , Aman Chadha , Vinija Jain , Suranjana Trivedy , Amitava Das

Vision-and-language navigation (VLN) is a multimodal task where an agent follows natural language instructions and navigates in visual environments. Multiple setups have been proposed, and researchers apply new model architectures or…

Computer Vision and Pattern Recognition · Computer Science 2022-05-05 Wanrong Zhu , Yuankai Qi , Pradyumna Narayana , Kazoo Sone , Sugato Basu , Xin Eric Wang , Qi Wu , Miguel Eckstein , William Yang Wang

In unknown cluttered and dynamic environments such as disaster scenes, mobile robots need to perform target-driven navigation in order to find people or objects of interest, while being solely guided by images of the targets. In this paper,…

Robotics · Computer Science 2024-07-09 Haitong Wang , Aaron Hao Tan , Goldie Nejat

Monocular visual odometry (VO) is an important task in robotics and computer vision. Thus far, how to build accurate and robust monocular VO systems that can work well in diverse scenarios remains largely unsolved. In this paper, we propose…

Computer Vision and Pattern Recognition · Computer Science 2022-04-05 Libo Sun , Wei Yin , Enze Xie , Zhengrong Li , Changming Sun , Chunhua Shen

Deep Reinforcement Learning (DRL) based navigation methods have demonstrated promising results for mobile robots, but suffer from limited action flexibility in confined spaces. Conventional DRL approaches predominantly learn forward-motion…

Robotics · Computer Science 2025-04-01 Shanze Wang , Mingao Tan , Zhibo Yang , Biao Huang , Xiaoyu Shen , Hailong Huang , Wei Zhang

We introduce a novel framework for training deep stereo networks effortlessly and without any ground-truth. By leveraging state-of-the-art neural rendering solutions, we generate stereo training data from image sequences collected with a…

Computer Vision and Pattern Recognition · Computer Science 2023-03-31 Fabio Tosi , Alessio Tonioni , Daniele De Gregorio , Matteo Poggi

We propose a depth map inference system from monocular videos based on a novel dataset for navigation that mimics aerial footage from gimbal stabilized monocular camera in rigid scenes. Unlike most navigation datasets, the lack of rotation…

Computer Vision and Pattern Recognition · Computer Science 2018-09-13 Clément Pinard , Laure Chevalley , Antoine Manzanera , David Filliat

A major focus of recent developments in stereo vision has been on how to obtain accurate dense disparity maps in passive stereo vision. Active vision systems enable more accurate estimations of dense disparity compared to passive stereo.…

Computer Vision and Pattern Recognition · Computer Science 2022-09-13 Laurent Valentin Jospin , Hamid Laga , Farid Boussaid , Mohammed Bennamoun

Blind people face a lot of problems in their daily routines. They have to struggle a lot just to do their day-to-day chores. In this paper, we have proposed a system with the objective to help the visually impaired by providing audio aid…

Computer Vision and Pattern Recognition · Computer Science 2019-11-21 Nikhil Thakurdesai , Anupam Tripathi , Dheeraj Butani , Smita Sankhe

The field of self-supervised monocular depth estimation has seen huge advancements in recent years. Most methods assume stereo data is available during training but usually under-utilize it and only treat it as a reference signal. We…

Computer Vision and Pattern Recognition · Computer Science 2019-05-02 Matan Goldman , Tal Hassner , Shai Avidan

Monocular vision-based Simultaneous Localization and Mapping (SLAM) is used for various purposes due to its advantages in cost, simple setup, as well as availability in the environments where navigation with satellites is not effective.…

Robotics · Computer Science 2018-10-03 Young-Hee Lee , Chen Zhu , Gabriele Giorgi , Christoph Günther

A major challenge in deploying the smallest of Micro Aerial Vehicle (MAV) platforms (< 100 g) is their inability to carry sensors that provide high-resolution metric depth information (e.g., LiDAR or stereo cameras). Current systems rely on…

Robotics · Computer Science 2023-11-27 Nathaniel Simon , Anirudha Majumdar

Monocular simultaneous localization and mapping (SLAM) algorithms estimate drone poses and build a 3D map using a single camera. Current algorithms include sparse methods that lack detailed geometry, while learning-driven approaches produce…

Robotics · Computer Science 2025-11-25 Jeryes Danial , Yosi Ben Asher , Itzik Klein

Accurate metric depth is critical for autonomous driving perception and simulation, yet current approaches struggle to achieve high metric accuracy, multi-view and temporal consistency, and cross-domain generalization. To address these…

Computer Vision and Pattern Recognition · Computer Science 2026-03-05 Qihao Sun , Jiarun Liu , Ziqian Ni , Jianyun Xu , Tao Xie , Lijun Zhao , Ruifeng Li , Sheng Yang

Depth from defocus (DfD) and stereo matching are two most studied passive depth sensing schemes. The techniques are essentially complementary: DfD can robustly handle repetitive textures that are problematic for stereo matching whereas…

Computer Vision and Pattern Recognition · Computer Science 2018-08-07 Zhang Chen , Xinqing Guo , Siyuan Li , Xuan Cao , Jingyi Yu

Estimating precise metric depth and scene reconstruction from monocular endoscopy is a fundamental task for surgical navigation in robotic surgery. However, traditional stereo matching adopts binocular images to perceive the depth…

Robotics · Computer Science 2022-11-29 Ruofeng Wei , Bin Li , Hangjie Mo , Fangxun Zhong , Yonghao Long , Qi Dou , Yun-Hui Liu , Dong Sun

As a fundamental vision task, stereo matching has made remarkable progress. While recent iterative optimization-based methods have achieved promising performance, their feature extraction capabilities still have room for improvement.…

Computer Vision and Pattern Recognition · Computer Science 2024-12-16 Jingyi Zhou , Haoyu Zhang , Jiakang Yuan , Peng Ye , Tao Chen , Hao Jiang , Meiya Chen , Yangyang Zhang

In this thesis, we leverage monocular cameras on aerial robots to predict depth and semantic maps in low-altitude unstructured environments. We propose a joint deep-learning architecture, named Co-SemDepth, that can perform the two tasks…

Computer Vision and Pattern Recognition · Computer Science 2026-03-17 Yara AlaaEldin

Humans are able to localize objects in the environment using both visual and auditory cues, integrating information from multiple modalities into a common reference frame. We introduce a system that can leverage unlabeled audio-visual data…

Computer Vision and Pattern Recognition · Computer Science 2019-10-28 Chuang Gan , Hang Zhao , Peihao Chen , David Cox , Antonio Torralba