中文
相关论文

相关论文: VOCAL: Visual Odometry via ContrAstive Learning

200 篇论文

Dynamic environments such as urban areas are still challenging for popular visual-inertial odometry (VIO) algorithms. Existing datasets typically fail to capture the dynamic nature of these environments, therefore making it difficult to…

机器人学 · 计算机科学 2021-02-12 Koji Minoda , Fabian Schilling , Valentin Wüest , Dario Floreano , Takehisa Yairi

Hybrid pipelines that combine deep learning with classical optimization have established themselves as the dominant approach to visual odometry (VO). By integrating neural network predictions with bundle adjustment, these models estimate…

计算机视觉与模式识别 · 计算机科学 2026-03-10 Vlardimir Yugay , Duy-Kien Nguyen , Theo Gevers , Cees G. M. Snoek , Martin R. Oswald

3D object detection plays a crucial role in autonomous systems, yet existing methods are limited by closed-set assumptions and struggle to recognize novel objects and their attributes in real-world scenarios. We propose OVODA, a novel…

计算机视觉与模式识别 · 计算机科学 2025-08-26 Xinhao Xiang , Kuan-Chuan Peng , Suhas Lohit , Michael J. Jones , Jiawei Zhang

The requirement for expert annotations limits the effectiveness of deep learning for medical image analysis. Although 3D self-supervised methods like volume contrast learning (VoCo) are powerful and partially address the labeling scarcity…

计算机视觉与模式识别 · 计算机科学 2026-01-22 Po-Kai Chiu , Hung-Hsuan Chen

Open-vocabulary 3D object detection for autonomous driving aims to detect novel objects beyond the predefined training label sets in point cloud scenes. Existing approaches achieve this by connecting traditional 3D object detectors with…

计算机视觉与模式识别 · 计算机科学 2025-03-11 Adrian Chow , Evelien Riddell , Yimu Wang , Sean Sedwards , Krzysztof Czarnecki

Multi-view geometry-based methods dominate the last few decades in monocular Visual Odometry for their superior performance, while they have been vulnerable to dynamic and low-texture scenes. More importantly, monocular methods suffer from…

计算机视觉与模式识别 · 计算机科学 2021-03-02 Huangying Zhan , Chamara Saroj Weerasekera , Jia-Wang Bian , Ravi Garg , Ian Reid

Visual-Inertial odometry (VIO) is the process of estimating the state (pose and velocity) of an agent (e.g., an aerial robot) by using only the input of one or more cameras plus one or more Inertial Measurement Units (IMUs) attached to it.…

机器人学 · 计算机科学 2019-06-17 Davide Scaramuzza , Zichao Zhang

Although cluttered indoor scenes have a lot of useful high-level semantic information which can be used for mapping and localization, most Visual Odometry (VO) algorithms rely on the usage of geometric features such as points, lines and…

计算机视觉与模式识别 · 计算机科学 2018-03-02 Huai-Jen Liang , Nitin J. Sanket , Cornelia Fermüller , Yiannis Aloimonos

Visual-Inertial Odometry (VIO) is a critical component for robust ego-motion estimation, enabling foundational capabilities such as autonomous navigation in robotics and real-time 6-DoF tracking for augmented reality. Existing methods face…

机器人学 · 计算机科学 2026-03-18 Feiyang Pan , Shenghe Zheng , Chunyan Yin , Guangbin Dou

Autonomous robots often rely on monocular cameras for odometry estimation and navigation. However, the scale ambiguity problem presents a critical barrier to effective monocular visual odometry. In this paper, we present CodedVO, a novel…

机器人学 · 计算机科学 2024-07-26 Sachin Shah , Naitri Rajyaguru , Chahat Deep Singh , Christopher Metzler , Yiannis Aloimonos

Odometry is of key importance for localization in the absence of a map. There is considerable work in the area of visual odometry (VO), and recent advances in deep learning have brought novel approaches to VO, which directly learn salient…

计算机视觉与模式识别 · 计算机科学 2020-03-06 Wei Wang , Muhamad Risqi U. Saputra , Peijun Zhao , Pedro Gusmao , Bo Yang , Changhao Chen , Andrew Markham , Niki Trigoni

Making multi-camera visual SLAM systems easier to set up and more robust to the environment is attractive for vision robots. Existing monocular and binocular vision SLAM systems have narrow sensing Field-of-View (FoV), resulting in…

机器人学 · 计算机科学 2025-03-26 Huai Yu , Junhao Wang , Yao He , Wen Yang , Gui-Song Xia

Recent work in visual representation learning for robotics demonstrates the viability of learning from large video datasets of humans performing everyday tasks. Leveraging methods such as masked autoencoding and contrastive learning, these…

机器人学 · 计算机科学 2023-02-27 Siddharth Karamcheti , Suraj Nair , Annie S. Chen , Thomas Kollar , Chelsea Finn , Dorsa Sadigh , Percy Liang

Underwater visual localization remains challenging due to wavelength-dependent attenuation, poor texture, and non-Gaussian sensor noise. We introduce MARVO, a physics-aware, learning-integrated odometry framework that fuses underwater image…

机器人学 · 计算机科学 2025-12-01 Sacchin Sundar , Atman Kikani , Aaliya Alam , Sumukh Shrote , A. Nayeemulla Khan , A. Shahina

Event-based cameras are bio-inspired sensors with pixels that independently and asynchronously respond to brightness changes at microsecond resolution, offering the potential to handle state estimation tasks involving motion blur and high…

机器人学 · 计算机科学 2025-09-11 Sheng Zhong , Junkai Niu , Yi Zhou

Existing image-text modality alignment in Vision Language Models (VLMs) treats each text token equally in an autoregressive manner. Despite being simple and effective, this method results in sub-optimal cross-modal alignment by…

计算机视觉与模式识别 · 计算机科学 2024-11-06 Xin Xiao , Bohong Wu , Jiacong Wang , Chunyuan Li , Xun Zhou , Haoyuan Guo

Visual-inertial odometry (VIO) is widely used in various fields, such as robots, drones, and autonomous vehicles. However, real-world scenes often feature dynamic objects, compromising the accuracy of VIO. The diversity and partial…

计算机视觉与模式识别 · 计算机科学 2025-03-04 Rui Zhou , Jingbin Liu , Junbin Xie , Jianyu Zhang , Yingze Hu , Jiele Zhao

This paper proposes an efficient and probabilistic adaptive voxel mapping method for LiDAR odometry. The map is a collection of voxels; each contains one plane (or edge) feature that enables the probabilistic representation of the…

机器人学 · 计算机科学 2022-07-11 Chongjian Yuan , Wei xu , Xiyuan Liu , Xiaoping Hong , Fu Zhang

This paper presents ViTOC (Vision Transformer and Object-aware Captioner), a novel vision-language model for image captioning that addresses the challenges of accuracy and diversity in generated descriptions. Unlike conventional approaches,…

计算机视觉与模式识别 · 计算机科学 2025-04-18 Feiyang Huang

Object detection (OD) in computer vision has made significant progress in recent years, transitioning from closed-set labels to open-vocabulary detection (OVD) based on large-scale vision-language pre-training (VLP). However, current…

计算机视觉与模式识别 · 计算机科学 2023-12-19 Yiyang Yao , Peng Liu , Tiancheng Zhao , Qianqian Zhang , Jiajia Liao , Chunxin Fang , Kyusong Lee , Qing Wang