中文
相关论文

相关论文: MVL-Loc: Leveraging Vision-Language Model for Gene…

200 篇论文

Absolute Visual Localization (AVL) enables an Unmanned Aerial Vehicle (UAV) to determine its position in GNSS-denied environments by establishing geometric relationships between UAV images and geo-tagged reference maps. While many previous…

计算机视觉与模式识别 · 计算机科学 2026-04-14 Yibin Ye , Xichao Teng , Shuo Chen , Leqi Liu , Kun Wang , Xiaokai Song , Zhang Li

Most existing approaches for visual localization either need a detailed 3D model of the environment or, in the case of learning-based methods, must be retrained for each new scene. This can either be very expensive or simply impossible for…

机器人学 · 计算机科学 2021-06-22 Dominik Winkelbauer , Maximilian Denninger , Rudolph Triebel

The monocular visual-inertial system (VINS), which consists one camera and one low-cost inertial measurement unit (IMU), is a popular approach to achieve accurate 6-DOF state estimation. However, such locally accurate visual-inertial…

计算机视觉与模式识别 · 计算机科学 2018-03-06 Tong Qin , Perliang Li , Shaojie Shen

Vision-Language Models (VLMs) have shown remarkable capabilities across diverse visual tasks, including image recognition, video understanding, and Visual Question Answering (VQA) when explicitly trained for these tasks. Despite these…

Beyond novel view synthesis, Neural Radiance Fields are useful for applications that interact with the real world. In this paper, we use them as an implicit map of a given scene and propose a camera relocalization algorithm tailored for…

计算机视觉与模式识别 · 计算机科学 2023-08-23 Arthur Moreau , Nathan Piasco , Moussab Bennehar , Dzmitry Tsishkou , Bogdan Stanciulescu , Arnaud de La Fortelle

In this paper, we present a tightly-coupled visual-inertial object-level multi-instance dynamic SLAM system. Even in extremely dynamic scenes, it can robustly optimise for the camera pose, velocity, IMU biases and build a dense 3D…

机器人学 · 计算机科学 2022-08-09 Yifei Ren , Binbin Xu , Christopher L. Choi , Stefan Leutenegger

Maps are a key component in image-based camera localization and visual SLAM systems: they are used to establish geometric constraints between images, correct drift in relative pose estimation, and relocalize cameras after lost tracking. The…

计算机视觉与模式识别 · 计算机科学 2018-04-03 Samarth Brahmbhatt , Jinwei Gu , Kihwan Kim , James Hays , Jan Kautz

Absolute camera pose regressors estimate the position and orientation of a camera given the captured image alone. Typically, a convolutional backbone with a multi-layer perceptron (MLP) head is trained using images and pose labels to embed…

计算机视觉与模式识别 · 计算机科学 2023-08-24 Yoli Shavit , Ron Ferens , Yosi Keller

Relocalization is the basis of map-based localization algorithms. Camera and LiDAR map-based methods are pervasive since their robustness under different scenarios. Generally, mapping and localization using the same sensor have better…

计算机视觉与模式识别 · 计算机科学 2023-05-16 Shuhang Tan , Hengyu Liu , Zhiling Wang

In this work, we tackle the problem of active camera localization, which controls the camera movements actively to achieve an accurate camera pose. The past solutions are mostly based on Markov Localization, which reduces the position-wise…

计算机视觉与模式识别 · 计算机科学 2022-07-21 Qihang Fang , Yingda Yin , Qingnan Fan , Fei Xia , Siyan Dong , Sheng Wang , Jue Wang , Leonidas Guibas , Baoquan Chen

The Simultaneous Localization and Mapping (SLAM) problem addresses the possibility of a robot to localize itself in an unknown environment and simultaneously build a consistent map of this environment. Recently, cameras have been…

计算机视觉与模式识别 · 计算机科学 2021-06-02 Hudson M. S. Bruno , Esther L. Colombini

We propose UnLoc, an efficient data-driven solution for sequential camera localization within floorplans. Floorplan data is readily available, long-term persistent, and robust to changes in visual appearance. We address key limitations of…

计算机视觉与模式识别 · 计算机科学 2026-03-17 Matthias Wüest , Francis Engelmann , Ondrej Miksik , Marc Pollefeys , Daniel Barath

Multimodal large language models (MLLMs), such as GPT-4o, Gemini, LLaVA, and Flamingo, have made significant progress in integrating visual and textual modalities, excelling in tasks like visual question answering (VQA), image captioning,…

计算机视觉与模式识别 · 计算机科学 2024-12-31 Junxiao Xue , Quan Deng , Fei Yu , Yanhao Wang , Jun Wang , Yuehua Li

Robots need the capability of placing objects in arbitrary, specific poses to rearrange the world and achieve various valuable tasks. Object reorientation plays a crucial role in this as objects may not initially be oriented such that the…

机器人学 · 计算机科学 2022-02-23 Kentaro Wada , Stephen James , Andrew J. Davison

Visual localization is a key step in many robotics pipelines, allowing the robot to (approximately) determine its position and orientation in the world. An efficient and scalable approach to visual localization is to use image retrieval…

计算机视觉与模式识别 · 计算机科学 2019-03-05 Asha Anoosheh , Torsten Sattler , Radu Timofte , Marc Pollefeys , Luc Van Gool

Visual localization is a fundamental task that regresses the 6 Degree Of Freedom (6DoF) poses with image features in order to serve the high precision localization requests in many robotics applications. Degenerate conditions like motion…

计算机视觉与模式识别 · 计算机科学 2022-10-19 Yuchen Yang , Xudong Zhang , Shuang Gao , Jixiang Wan , Yishan Ping , Yuyue Liu , Jijunnan Li , Yandong Guo

Visual localization is the task of estimating a 6-DoF camera pose of a query image within a provided 3D reference map. Thanks to recent advances in various 3D sensors, 3D point clouds are becoming a more accurate and affordable option for…

计算机视觉与模式识别 · 计算机科学 2023-09-15 Minjung Kim , Junseo Koo , Gunhee Kim

Camera localization is a fundamental and key component of autonomous driving vehicles and mobile robots to localize themselves globally for further environment perception, path planning and motion control. Recently end-to-end approaches…

计算机视觉与模式识别 · 计算机科学 2020-05-14 Mi Tian , Qiong Nie , Hao Shen

As the number of installed cameras grows, so do the compute resources required to process and analyze all the images captured by these cameras. Video analytics enables new use cases, such as smart cities or autonomous driving. At the same…

计算机视觉与模式识别 · 计算机科学 2021-12-01 Daniel Rivas , Francesc Guim , Jordà Polo , David Carrera

This paper proposes SOLVR, a unified pipeline for learning based LiDAR-Visual re-localisation which performs place recognition and 6-DoF registration across sensor modalities. We propose a strategy to align the input sensor modalities by…

计算机视觉与模式识别 · 计算机科学 2025-02-07 Joshua Knights , Sebastián Barbas Laina , Peyman Moghadam , Stefan Leutenegger