English
Related papers

Related papers: Unsupervised Visual Odometry and Action Integratio…

200 papers

Accurate relative pose is one of the key components in visual odometry (VO) and simultaneous localization and mapping (SLAM). Recently, the self-supervised learning framework that jointly optimizes the relative pose and target image depth…

Computer Vision and Pattern Recognition · Computer Science 2019-02-26 Tianwei Shen , Zixin Luo , Lei Zhou , Hanyu Deng , Runze Zhang , Tian Fang , Long Quan

Mobile robots exploring indoor environments increasingly rely on vision-language models to perceive high-level semantic cues in camera images, such as object categories. Such models offer the potential to substantially advance robot…

Robotics · Computer Science 2025-10-09 Utkarsh Bajpai , Julius Rückin , Cyrill Stachniss , Marija Popović

Object goal navigation aims to navigate an agent to locations of a given object category in unseen environments. Classical methods explicitly build maps of environments and require extensive engineering while lacking semantic information…

Computer Vision and Pattern Recognition · Computer Science 2023-08-11 Shizhe Chen , Thomas Chabal , Ivan Laptev , Cordelia Schmid

Monocular visual navigation methods have seen significant advances in the last decade, recently producing several real-time solutions for autonomously navigating small unmanned aircraft systems without relying on GPS. This is critical for…

Computer Vision and Pattern Recognition · Computer Science 2020-09-24 Kyung Kim , Robert C. Leishman , Scott L. Nykl

Generally, high-level features provide more geometrical information compared to point features, which can be exploited to further constrain motions. Planes are commonplace in man-made environments, offering an active means to reduce drift,…

Robotics · Computer Science 2025-05-20 Yidi Zhang , Fulin Tang , Zewen Xu , Yihong Wu , Pengju Ma

Robust and accurate localization for Unmanned Aerial Vehicles (UAVs) is an essential capability to achieve autonomous, long-range flights. Current methods either rely heavily on GNSS, face limitations in visual-based localization due to…

Robotics · Computer Science 2023-10-26 Yao He , Ivan Cisneros , Nikhil Keetha , Jay Patrikar , Zelin Ye , Ian Higgins , Yaoyu Hu , Parv Kapoor , Sebastian Scherer

We present ConVOI, a novel method for autonomous robot navigation in real-world indoor and outdoor environments using Vision Language Models (VLMs). We employ VLMs in two ways: first, we leverage their zero-shot image classification…

The amount of texture can be rich or deficient depending on the objects and the structures of the building. The conventional mono visual-initial navigation system (VINS)-based localization techniques perform well in environments where…

Robotics · Computer Science 2021-01-01 KwangYik Jung , YeEun Kim , HyunJun Lim , Hyun Myung

We study the challenging problem of releasing a robot in a previously unseen environment, and having it follow unconstrained natural language navigation instructions. Recent work on the task of Vision-and-Language Navigation (VLN) has…

Computer Vision and Pattern Recognition · Computer Science 2020-11-10 Peter Anderson , Ayush Shrivastava , Joanne Truong , Arjun Majumdar , Devi Parikh , Dhruv Batra , Stefan Lee

Unsupervised localization and segmentation are long-standing robot vision challenges that describe the critical ability for an autonomous robot to learn to decompose images into individual objects without labeled data. These tasks are…

Computer Vision and Pattern Recognition · Computer Science 2023-07-26 Xinyu Zhang , Abdeslam Boularias

Unsupervised Learning based monocular visual odometry (VO) has lately drawn significant attention for its potential in label-free leaning ability and robustness to camera parameters and environmental variations. However, partially due to…

Computer Vision and Pattern Recognition · Computer Science 2019-03-18 Yang Li , Yoshitaka Ushiku , Tatsuya Harada

We propose a self-supervised learning framework that uses unlabeled monocular video sequences to generate large-scale supervision for training a Visual Odometry (VO) frontend, a network which computes pointwise data associations across…

Computer Vision and Pattern Recognition · Computer Science 2018-12-11 Daniel DeTone , Tomasz Malisiewicz , Andrew Rabinovich

Dynamic environments such as urban areas are still challenging for popular visual-inertial odometry (VIO) algorithms. Existing datasets typically fail to capture the dynamic nature of these environments, therefore making it difficult to…

Robotics · Computer Science 2021-02-12 Koji Minoda , Fabian Schilling , Valentin Wüest , Dario Floreano , Takehisa Yairi

We propose a novel monocular visual odometry (VO) system called UnDeepVO in this paper. UnDeepVO is able to estimate the 6-DoF pose of a monocular camera and the depth of its view by using deep neural networks. There are two salient…

Computer Vision and Pattern Recognition · Computer Science 2018-02-22 Ruihao Li , Sen Wang , Zhiqiang Long , Dongbing Gu

Navigating autonomous underwater vehicles (AUVs) in unknown environments is significantly challenging due to poor visibility, weak signal transmission, and dynamic water currents. These factors pose challenges in accurate global…

Pretrained video generation models provide strong priors for robot control, but existing unified world action models still struggle to decode reliable actions without substantial robot-specific training. We attribute this limitation to a…

Robotics · Computer Science 2026-04-14 Liaoyuan Fan , Zetian Xu , Chen Cao , Wenyao Zhang , Mingqi Yuan , Jiayu Chen

Traditional Visual Odometry (VO) and Visual Inertial Odometry (VIO) methods rely on a 'pose-centric' paradigm, which computes absolute camera poses from the local map thus requires large-scale landmark maintenance and continuous map…

Robotics · Computer Science 2025-11-13 Sangheon Yang , Yeongin Yoon , Hong Mo Jung , Jongwoo Lim

Audio-visual embodied navigation, as a hot research topic, aims training a robot to reach an audio target using egocentric visual (from the sensors mounted on the robot) and audio (emitted from the target) input. The audio-visual…

Sound · Computer Science 2022-10-06 Yinfeng Yu , Lele Cao , Fuchun Sun , Xiaohong Liu , Liejun Wang

In the context of autonomous navigation of terrestrial robots, the creation of realistic models for agent dynamics and sensing is a widespread habit in the robotics literature and in commercial applications, where they are used for model…

The advances in deep reinforcement learning recently revived interest in data-driven learning based approaches to navigation. In this paper we propose to learn viewpoint invariant and target invariant visual servoing for local mobile robot…

Computer Vision and Pattern Recognition · Computer Science 2020-03-06 Yimeng Li , Jana Kosecka