中文
相关论文

相关论文: GaGA: Towards Interactive Global Geolocation Assis…

200 篇论文

Image geolocalization, the task of identifying the geographic location depicted in an image, is important for applications in crisis response, digital forensics, and location-based intelligence. While recent advances in large language…

计算机视觉与模式识别 · 计算机科学 2026-05-19 Lingyao Li , Runlong Yu , Qikai Hu , Bowei Li , Min Deng , Yang Zhou , Xiaowei Jia

Due to the advances in mobile computing and multimedia techniques, there are vast amount of multimedia data with geographical information collected in multifarious applications. In this paper, we propose a novel type of image search named…

多媒体 · 计算机科学 2018-06-05 Jun Long , Lei Zhu , Chengyuan Zhang , Zhan Yang , Yunwu Lin , Ruipeng Chen

Cross-view geo-localisation identifies coarse geographical position of an automated vehicle by matching a ground-level image to a geo-tagged satellite image from a database. Despite the advancements in Cross-view geo-localisation,…

计算机视觉与模式识别 · 计算机科学 2025-05-21 Barkin Dagda , Muhammad Awais , Saber Fallah

Large language models (LLMs) are being used in data science code generation tasks, but they often struggle with complex sequential tasks, leading to logical errors. Their application to geospatial data processing is particularly challenging…

计算机与社会 · 计算机科学 2024-10-28 Yuxing Chen , Weijie Wang , Sylvain Lobry , Camille Kurtz

Planet-scale photo geolocalization is the complex task of estimating the location depicted in an image solely based on its visual content. Due to the success of convolutional neural networks (CNNs), current approaches achieve super-human…

计算机视觉与模式识别 · 计算机科学 2021-10-22 Jonas Theiner , Eric Mueller-Budack , Ralph Ewerth

The application of machine learning (ML) in a range of geospatial tasks is increasingly common but often relies on globally available covariates such as satellite imagery that can either be expensive or lack predictive power. Here we…

计算与语言 · 计算机科学 2024-02-27 Rohin Manvi , Samar Khanna , Gengchen Mai , Marshall Burke , David Lobell , Stefano Ermon

This paper presents GeoAgent, a model capable of reasoning closely with humans and deriving fine-grained address conclusions. Previous RL-based methods have achieved breakthroughs in performance and interpretability but still remain…

人工智能 · 计算机科学 2026-02-16 Modi Jin , Yiming Zhang , Boyuan Sun , Dingwen Zhang , MingMing Cheng , Qibin Hou

The image geolocalization task aims to predict the location where an image was taken anywhere on Earth using visual clues. Existing large vision-language model (LVLM) approaches leverage world knowledge, chain-of-thought reasoning, and…

计算机视觉与模式识别 · 计算机科学 2026-01-12 Yuxiang Ji , Yong Wang , Ziyu Ma , Yiming Hu , Hailang Huang , Xuecai Hu , Guanhua Chen , Liaoni Wu , Xiangxiang Chu

The existing work in cross-view geo-localization is based on images where a ground panorama is matched to an aerial image. In this work, we focus on ground videos instead of images which provides additional contextual cues which are…

计算机视觉与模式识别 · 计算机科学 2022-07-07 Shruti Vyas , Chen Chen , Mubarak Shah

Cross-view geolocalization, a supplement or replacement for GPS, localizes an agent within a search area by matching images taken from a ground-view camera to overhead images taken from satellites or aircraft. Although the viewpoint…

机器人学 · 计算机科学 2023-05-19 Lena M. Downes , Ted J. Steiner , Rebecca L. Russell , Jonathan P. How

Image geolocalization, the task of determining an image's geographic origin, poses significant challenges, largely due to visual similarities across disparate locations and the large search space. To address these issues, we propose a…

计算机视觉与模式识别 · 计算机科学 2025-11-04 Narges Ghasemi , Amir Ziashahabi , Salman Avestimehr , Cyrus Shahabi

Graph Neural Networks (GNNs) have empowered the advance in graph-structured data analysis. Recently, the rise of Large Language Models (LLMs) like GPT-4 has heralded a new era in deep learning. However, their application to graph data poses…

机器学习 · 计算机科学 2024-04-12 Runjin Chen , Tong Zhao , Ajay Jaiswal , Neil Shah , Zhangyang Wang

Recent advances in Large Vision-Language Models (LVLMs) have significantly improve performance in image comprehension tasks, such as formatted charts and rich-content images. Yet, Graphical User Interface (GUI) pose a greater challenge due…

计算机视觉与模式识别 · 计算机科学 2024-11-05 Ziyang Meng , Yu Dai , Zezheng Gong , Shaoxiong Guo , Minglong Tang , Tongquan Wei

The prevalence of Vision-Language Models (VLMs) raises important questions about privacy in an era where visual information is increasingly available. While foundation VLMs demonstrate broad knowledge and learned capabilities, we…

计算机视觉与模式识别 · 计算机科学 2025-02-21 Neel Jay , Hieu Minh Nguyen , Trung Dung Hoang , Jacob Haimes

This paper proposes a novel method for vision-based metric cross-view geolocalization (CVGL) that matches the camera images captured from a ground-based vehicle with an aerial image to determine the vehicle's geo-pose. Since aerial images…

计算机视觉与模式识别 · 计算机科学 2023-05-18 Florian Fervers , Sebastian Bullinger , Christoph Bodensteiner , Michael Arens , Rainer Stiefelhagen

Estimating the location where an image was taken based solely on the contents of the image is a challenging task, even for humans, as properly labeling an image in such a fashion relies heavily on contextual information, and is not as…

计算机视觉与模式识别 · 计算机科学 2017-12-29 Jesse M. Johns , Jeremiah Rounds , Michael J. Henry

Cross-view geo-localization identifies the locations of street-view images by matching them with geo-tagged satellite images or OSM. However, most existing studies focus on image-to-image retrieval, with fewer addressing text-guided…

计算机视觉与模式识别 · 计算机科学 2025-04-02 Junyan Ye , Honglin Lin , Leyan Ou , Dairong Chen , Zihao Wang , Qi Zhu , Conghui He , Weijia Li

Image geolocation is a critical task in various image-understanding applications. However, existing methods often fail when analyzing challenging, in-the-wild images. Inspired by the exceptional background knowledge of multimodal language…

计算机视觉与模式识别 · 计算机科学 2024-06-03 Zhiqiang Wang , Dejia Xu , Rana Muhammad Shahroz Khan , Yanbin Lin , Zhiwen Fan , Xingquan Zhu

Vision-Language-Action (VLA) models achieve strong generalization in robotic manipulation but remain largely reactive and 2D-centric, making them unreliable in tasks that require precise 3D reasoning. We propose GeoPredict, a geometry-aware…

计算机视觉与模式识别 · 计算机科学 2026-04-08 Jingjing Qian , Boyao Han , Chen Shi , Lei Xiao , Long Yang , Shaoshuai Shi , Li Jiang

In geographic data videos, camera movements are frequently used and combined to present information from multiple perspectives. However, creating and editing camera movements requires significant time and professional skills. This work aims…

人机交互 · 计算机科学 2023-09-12 Wenchao Li , Zhan Wang , Yun Wang , Di Weng , Liwenhan Xie , Siming Chen , Haidong Zhang , Huamin Qu