中文
相关论文

相关论文: G^3: Geolocation via Guidebook Grounding

200 篇论文

For robots to understand human instructions and perform meaningful tasks in the near future, it is important to develop learned models that comprehend referential language to identify common objects in real-world 3D scenes. In this paper,…

机器人学 · 计算机科学 2021-11-08 Junha Roh , Karthik Desingh , Ali Farhadi , Dieter Fox

We study the task of locating a user in a mapped indoor environment using natural language queries and images from the environment. Building on recent pretrained vision-language models, we learn a similarity score between text descriptions…

计算机视觉与模式识别 · 计算机科学 2024-10-08 Seth Pate , Lawson L. S. Wong

Understanding human instructions is essential for enabling smooth human-robot interaction. In this work, we focus on object grounding, i.e., localizing an object of interest in a visual scene (e.g., an image) based on verbal human…

计算机视觉与模式识别 · 计算机科学 2025-12-01 Joel Alberto Santos , Zongwei Wu , Xavier Alameda-Pineda , Radu Timofte

3-Dimensional Embodied Reference Understanding (3D-ERU) combines a language description and an accompanying pointing gesture to identify the most relevant target object in a 3D scene. Although prior work has explored pure language-based 3D…

计算机视觉与模式识别 · 计算机科学 2025-04-15 Atharv Mahesh Mane , Dulanga Weerakoon , Vigneshwaran Subbaraju , Sougata Sen , Sanjay E. Sarma , Archan Misra

Vision-language models (VLMs) have shown a promising ability in image geolocation, but they still lack structured geographic reasoning and the capacity for autonomous self-evolution. Existing methods predominantly rely on implicit…

计算机视觉与模式识别 · 计算机科学 2026-05-14 Chenjie Yang , Yutian Jiang , Yutong Deng , Chenyu Wu

Language is highly structured, with syntactic and semantic structures, to some extent, agreed upon by speakers of the same language. With implicit or explicit awareness of such structures, humans can learn and use language efficiently and…

计算与语言 · 计算机科学 2024-10-23 Freda Shi

Image geolocalization, in which an AI model traditionally predicts the precise GPS coordinates of an image, is a challenging task with many downstream applications. However, the user cannot utilize the model to further their knowledge…

计算机视觉与模式识别 · 计算机科学 2025-09-04 Ron Campos , Ashmal Vayani , Parth Parag Kulkarni , Rohit Gupta , Aizan Zafar , Aritra Dutta , Mubarak Shah

From a single image, visual cues can help deduce intrinsic and extrinsic camera parameters like the focal length and the gravity direction. This single-image calibration can benefit various downstream applications like image editing and 3D…

计算机视觉与模式识别 · 计算机科学 2024-10-18 Alexander Veicht , Paul-Edouard Sarlin , Philipp Lindenberger , Marc Pollefeys

This paper presents a visual geo-localization system capable of determining the geographic locations of places (buildings and road intersections) from images without relying on GPS data. Our approach integrates three primary methods:…

计算机视觉与模式识别 · 计算机科学 2024-07-23 Rania Saoud , Slimane Larabi

Vision-language models (VLMs) have advanced rapidly, yet their capacity for image-grounded geolocation in open-world conditions, a task that is challenging and of demand in real life, has not been comprehensively evaluated. We present…

计算机视觉与模式识别 · 计算机科学 2025-10-14 Zhaofang Qian , Hardy Chen , Zeyu Wang , Li Zhang , Zijun Wang , Xiaoke Huang , Hui Liu , Xianfeng Tang , Zeyu Zheng , Haoqin Tu , Cihang Xie , Yuyin Zhou

Place recognition, the ability to identify previously visited locations, is critical for both biological navigation and autonomous systems. This review synthesizes findings from robotic systems, animal studies, and human research to explore…

机器人学 · 计算机科学 2025-11-19 Michael Milford , Tobias Fischer

Localization of a robotic system within a previously mapped environment is important for reducing estimation drift and for reusing previously built maps. Existing techniques for geometry-based localization have focused on the description of…

Understanding terrain topology at long-range is crucial for the success of off-road robotic missions, especially when navigating at high-speeds. LiDAR sensors, which are currently heavily relied upon for geometric mapping, provide sparse…

机器人学 · 计算机科学 2024-04-23 Chanyoung Chung , Georgios Georgakis , Patrick Spieler , Curtis Padgett , Ali Agha , Shehryar Khattak

Today's autonomous vehicles rely extensively on high-definition 3D maps to navigate the environment. While this approach works well when these maps are completely up-to-date, safe autonomous vehicles must be able to corroborate the map's…

计算机视觉与模式识别 · 计算机科学 2016-12-09 Ari Seff , Jianxiong Xiao

This work presents a method that is able to predict the geolocation of a street-view photo taken in the wild within a state-sized search region by matching against a database of aerial reference imagery. We partition the search region into…

计算机视觉与模式识别 · 计算机科学 2024-09-26 Florian Fervers , Sebastian Bullinger , Christoph Bodensteiner , Michael Arens , Rainer Stiefelhagen

Accurate metrical localization is one of the central challenges in mobile robotics. Many existing methods aim at localizing after building a map with the robot. In this paper, we present a novel approach that instead uses geotagged…

机器人学 · 计算机科学 2015-04-17 Pratik Agarwal , Wolfram Burgard , Luciano Spinello

Recently, language-guided global image editing draws increasing attention with growing application potentials. However, previous GAN-based methods are not only confined to domain-specific, low-resolution data but also lacking in…

计算机视觉与模式识别 · 计算机科学 2021-06-25 Jing Shi , Ning Xu , Yihang Xu , Trung Bui , Franck Dernoncourt , Chenliang Xu

The emergence of Vision-Language Models (VLMs) has introduced new paradigms for global image geo-localization through retrieval-augmented generation (RAG) and reasoning-driven inference. However, RAG methods are constrained by retrieval…

计算机视觉与模式识别 · 计算机科学 2026-03-17 Bo Yu , Fengze Yang , Yiming Liu , Chao Wang , Xuewen Luo , Taozhe Li , Ruimin Ke , Xiaofan Zhou , Chenxi Liu

Given a textual phrase and an image, the visual grounding problem is the task of locating the content of the image referenced by the sentence. It is a challenging task that has several real-world applications in human-computer interaction,…

计算机视觉与模式识别 · 计算机科学 2022-02-03 Davide Rigoni , Luciano Serafini , Alessandro Sperduti

Despite significant algorithmic advances in vision-based positioning, a comprehensive probabilistic framework to study its performance has remained unexplored. The main objective of this paper is to develop such a framework using ideas from…

信息论 · 计算机科学 2024-09-17 Haozhou Hu , Harpreet S. Dhillon , R. Michael Buehrer