English
Related papers

Related papers: G^3: Geolocation via Guidebook Grounding

200 papers

For robots to understand human instructions and perform meaningful tasks in the near future, it is important to develop learned models that comprehend referential language to identify common objects in real-world 3D scenes. In this paper,…

Robotics · Computer Science 2021-11-08 Junha Roh , Karthik Desingh , Ali Farhadi , Dieter Fox

We study the task of locating a user in a mapped indoor environment using natural language queries and images from the environment. Building on recent pretrained vision-language models, we learn a similarity score between text descriptions…

Computer Vision and Pattern Recognition · Computer Science 2024-10-08 Seth Pate , Lawson L. S. Wong

Understanding human instructions is essential for enabling smooth human-robot interaction. In this work, we focus on object grounding, i.e., localizing an object of interest in a visual scene (e.g., an image) based on verbal human…

Computer Vision and Pattern Recognition · Computer Science 2025-12-01 Joel Alberto Santos , Zongwei Wu , Xavier Alameda-Pineda , Radu Timofte

3-Dimensional Embodied Reference Understanding (3D-ERU) combines a language description and an accompanying pointing gesture to identify the most relevant target object in a 3D scene. Although prior work has explored pure language-based 3D…

Computer Vision and Pattern Recognition · Computer Science 2025-04-15 Atharv Mahesh Mane , Dulanga Weerakoon , Vigneshwaran Subbaraju , Sougata Sen , Sanjay E. Sarma , Archan Misra

Vision-language models (VLMs) have shown a promising ability in image geolocation, but they still lack structured geographic reasoning and the capacity for autonomous self-evolution. Existing methods predominantly rely on implicit…

Computer Vision and Pattern Recognition · Computer Science 2026-05-14 Chenjie Yang , Yutian Jiang , Yutong Deng , Chenyu Wu

Language is highly structured, with syntactic and semantic structures, to some extent, agreed upon by speakers of the same language. With implicit or explicit awareness of such structures, humans can learn and use language efficiently and…

Computation and Language · Computer Science 2024-10-23 Freda Shi

Image geolocalization, in which an AI model traditionally predicts the precise GPS coordinates of an image, is a challenging task with many downstream applications. However, the user cannot utilize the model to further their knowledge…

Computer Vision and Pattern Recognition · Computer Science 2025-09-04 Ron Campos , Ashmal Vayani , Parth Parag Kulkarni , Rohit Gupta , Aizan Zafar , Aritra Dutta , Mubarak Shah

From a single image, visual cues can help deduce intrinsic and extrinsic camera parameters like the focal length and the gravity direction. This single-image calibration can benefit various downstream applications like image editing and 3D…

Computer Vision and Pattern Recognition · Computer Science 2024-10-18 Alexander Veicht , Paul-Edouard Sarlin , Philipp Lindenberger , Marc Pollefeys

This paper presents a visual geo-localization system capable of determining the geographic locations of places (buildings and road intersections) from images without relying on GPS data. Our approach integrates three primary methods:…

Computer Vision and Pattern Recognition · Computer Science 2024-07-23 Rania Saoud , Slimane Larabi

Vision-language models (VLMs) have advanced rapidly, yet their capacity for image-grounded geolocation in open-world conditions, a task that is challenging and of demand in real life, has not been comprehensively evaluated. We present…

Computer Vision and Pattern Recognition · Computer Science 2025-10-14 Zhaofang Qian , Hardy Chen , Zeyu Wang , Li Zhang , Zijun Wang , Xiaoke Huang , Hui Liu , Xianfeng Tang , Zeyu Zheng , Haoqin Tu , Cihang Xie , Yuyin Zhou

Place recognition, the ability to identify previously visited locations, is critical for both biological navigation and autonomous systems. This review synthesizes findings from robotic systems, animal studies, and human research to explore…

Robotics · Computer Science 2025-11-19 Michael Milford , Tobias Fischer

Localization of a robotic system within a previously mapped environment is important for reducing estimation drift and for reusing previously built maps. Existing techniques for geometry-based localization have focused on the description of…

Understanding terrain topology at long-range is crucial for the success of off-road robotic missions, especially when navigating at high-speeds. LiDAR sensors, which are currently heavily relied upon for geometric mapping, provide sparse…

Today's autonomous vehicles rely extensively on high-definition 3D maps to navigate the environment. While this approach works well when these maps are completely up-to-date, safe autonomous vehicles must be able to corroborate the map's…

Computer Vision and Pattern Recognition · Computer Science 2016-12-09 Ari Seff , Jianxiong Xiao

This work presents a method that is able to predict the geolocation of a street-view photo taken in the wild within a state-sized search region by matching against a database of aerial reference imagery. We partition the search region into…

Computer Vision and Pattern Recognition · Computer Science 2024-09-26 Florian Fervers , Sebastian Bullinger , Christoph Bodensteiner , Michael Arens , Rainer Stiefelhagen

Accurate metrical localization is one of the central challenges in mobile robotics. Many existing methods aim at localizing after building a map with the robot. In this paper, we present a novel approach that instead uses geotagged…

Robotics · Computer Science 2015-04-17 Pratik Agarwal , Wolfram Burgard , Luciano Spinello

Recently, language-guided global image editing draws increasing attention with growing application potentials. However, previous GAN-based methods are not only confined to domain-specific, low-resolution data but also lacking in…

Computer Vision and Pattern Recognition · Computer Science 2021-06-25 Jing Shi , Ning Xu , Yihang Xu , Trung Bui , Franck Dernoncourt , Chenliang Xu

The emergence of Vision-Language Models (VLMs) has introduced new paradigms for global image geo-localization through retrieval-augmented generation (RAG) and reasoning-driven inference. However, RAG methods are constrained by retrieval…

Computer Vision and Pattern Recognition · Computer Science 2026-03-17 Bo Yu , Fengze Yang , Yiming Liu , Chao Wang , Xuewen Luo , Taozhe Li , Ruimin Ke , Xiaofan Zhou , Chenxi Liu

Given a textual phrase and an image, the visual grounding problem is the task of locating the content of the image referenced by the sentence. It is a challenging task that has several real-world applications in human-computer interaction,…

Computer Vision and Pattern Recognition · Computer Science 2022-02-03 Davide Rigoni , Luciano Serafini , Alessandro Sperduti

Despite significant algorithmic advances in vision-based positioning, a comprehensive probabilistic framework to study its performance has remained unexplored. The main objective of this paper is to develop such a framework using ideas from…

Information Theory · Computer Science 2024-09-17 Haozhou Hu , Harpreet S. Dhillon , R. Michael Buehrer