中文
相关论文

相关论文: G^3: Geolocation via Guidebook Grounding

200 篇论文

Vision based localization is the problem of inferring the pose of the camera given a single image. One solution to this problem is to learn a deep neural network to infer the pose of a query image after learning on a dataset of images with…

机器学习 · 计算机科学 2019-11-11 Carlos Lassance , Yasir Latif , Ravi Garg , Vincent Gripon , Ian Reid

This paper presents Vision-Language Global Localization (VLG-Loc), a novel global localization method that uses human-readable labeled footprint maps containing only names and areas of distinctive visual landmarks in an environment. While…

机器人学 · 计算机科学 2025-12-19 Mizuho Aoki , Kohei Honda , Yasuhiro Yoshimura , Takeshi Ishita , Ryo Yonetani

Multimodal intelligence development recently show strong progress in visual understanding and high level reasoning. Though, most reasoning system still reply on textual information as the main medium for inference. This limit their…

机器学习 · 计算机科学 2026-01-01 Soham Pahari , M. Srinivas

Current research on agentic visual reasoning enables deep multimodal understanding but primarily focuses on image manipulation tools, leaving a gap toward more general-purpose agentic models. In this work, we revisit the geolocalization…

计算机视觉与模式识别 · 计算机科学 2025-12-19 Yikun Wang , Zuyan Liu , Ziyi Wang , Han Hu , Pengfei Liu , Yongming Rao

We study the image-based geolocalization problem, aiming to localize ground-view query images on cartographic maps. Current methods often utilize cross-view localization techniques to match ground-view query images with 2D maps. However,…

计算机视觉与模式识别 · 计算机科学 2023-11-06 Mengjie Zhou , Liu Liu , Yiran Zhong , Andrew Calway

The application of machine learning (ML) in a range of geospatial tasks is increasingly common but often relies on globally available covariates such as satellite imagery that can either be expensive or lack predictive power. Here we…

计算与语言 · 计算机科学 2024-02-27 Rohin Manvi , Samar Khanna , Gengchen Mai , Marshall Burke , David Lobell , Stefano Ermon

Sensemaking using automatically extracted information from text is a challenging problem. In this paper, we address a specific type of information extraction, namely extracting information related to descriptions of movement. Aggregating…

人机交互 · 计算机科学 2022-04-21 Scott Pezanowski , Prasenjit Mitra , Alan M. MacEachren

The task of textual geolocation - retrieving the coordinates of a place based on a free-form language description - calls for not only grounding but also natural language understanding and geospatial reasoning. Even though there are quite a…

计算与语言 · 计算机科学 2023-07-04 Tzuf Paz-Argaman , Tal Bauman , Itai Mondshine , Itzhak Omer , Sagi Dalyot , Reut Tsarfaty

Many task domains require robots to interpret and act upon natural language commands which are given by people and which refer to the robot's physical surroundings. Such interpretation is known variously as the symbol grounding problem,…

Georeferencing text documents has typically relied on either gazetteer-based methods to assign geographic coordinates to place names, or on language modelling approaches that associate textual terms with geographic locations. However, many…

人工智能 · 计算机科学 2026-01-26 Aneesha Fernando , Surangika Ranathunga , Kristin Stock , Raj Prasanna , Christopher B. Jones

When driving, people make decisions based on current traffic as well as their desired route. They have a mental map of known routes and are often able to navigate without needing directions. Current self-driving models improve their…

计算机视觉与模式识别 · 计算机科学 2019-10-08 Iulia Paraicu , Marius Leordeanu

Web search engines have long served as indispensable tools for information retrieval; user behavior and query formulation strategies have been well studied. The introduction of search engines powered by large language models (LLMs)…

信息检索 · 计算机科学 2024-01-19 Albatool Wazzan , Stephen MacNeil , Richard Souvenir

Image geo-localization aims to determine where a photograph was taken, a task that often requires more than recognizing visible landmarks. Human experts typically solve it through an iterative workflow: they inspect informative regions,…

计算机视觉与模式识别 · 计算机科学 2026-05-27 Yong Li , Furong Jia , Dacheng Yin , Kang Rong , Fengyun Rao , Jing Lyu , Fan Zhang

In this paper, we propose to go beyond the well-established approach to vision-based localization that relies on visual descriptor matching between a query image and a 3D point cloud. While matching keypoints via visual descriptors makes…

计算机视觉与模式识别 · 计算机科学 2022-08-02 Qunjie Zhou , Sérgio Agostinho , Aljosa Osep , Laura Leal-Taixé

How does a person work out their location using a floorplan? It is probably safe to say that we do not explicitly measure depths to every visible surface and try to match them against different pose estimates in the floorplan. And yet, this…

机器人学 · 计算机科学 2018-05-03 Oscar Mendez , Simon Hadfield , Nicolas Pugeault , Richard Bowden

Worldwide image geolocalization, which aims to predict the GPS coordinates of any image on Earth, remains challenging due to global visual diversity. Recent generative approaches based on Retrieval-Augmented Generation (RAG) and Large…

信息检索 · 计算机科学 2026-04-29 Tung-Duong Le-Duc , Hoang-Quoc Nguyen-Son , Minh-Son Dao

In this work we present a novel approach to joint semantic localisation and scene understanding. Our work is motivated by the need for localisation algorithms which not only predict 6-DoF camera pose but also simultaneously recognise…

计算机视觉与模式识别 · 计算机科学 2019-09-24 Ignas Budvytis , Marvin Teichmann , Tomas Vojir , Roberto Cipolla

Graphical User Interface (GUI) grounding is commonly framed as a coordinate prediction task -- given a natural language instruction, generate on-screen coordinates for actions such as clicks and keystrokes. However, recent Vision Language…

A vast amount of location information exists in unstructured texts, such as social media posts, news stories, scientific articles, web pages, travel blogs, and historical archives. Geoparsing refers to the process of recognizing location…

计算与语言 · 计算机科学 2022-07-06 Xuke Hu , Zhiyong Zhou , Hao Li , Yingjie Hu , Fuqiang Gu , Jens Kersten , Hongchao Fan , Friederike Klan

While pretrained language models (PLMs) have been shown to possess a plethora of linguistic knowledge, the existing body of research has largely neglected extralinguistic knowledge, which is generally difficult to obtain by pretraining on…

计算与语言 · 计算机科学 2024-01-30 Valentin Hofmann , Goran Glavaš , Nikola Ljubešić , Janet B. Pierrehumbert , Hinrich Schütze