English
Related papers

Related papers: Street-Level Geolocalization Using Multimodal Larg…

200 papers

Deep learning has significantly advanced building segmentation in remote sensing, yet models struggle to generalize on data of diverse geographic regions due to variations in city layouts and the distribution of building types, sizes and…

Computer Vision and Pattern Recognition · Computer Science 2025-07-23 Shuang Song , Yang Tang , Rongjun Qin

Geolocation is a fundamental component of route planning and navigation for unmanned vehicles, but GNSS-based geolocation fails under denial-of-service conditions. Cross-view geo-localization (CVGL), which aims to estimate the geographical…

Computer Vision and Pattern Recognition · Computer Science 2023-04-21 Jianwei Zhao , Qiang Zhai , Pengbo Zhao , Rui Huang , Hong Cheng

Images shared on social media often expose geographic cues. While early geolocation methods required expert effort and lacked generalization, the rise of Large Vision Language Models (LVLMs) now enables accurate geolocation even for…

Cryptography and Security · Computer Science 2025-12-01 Xinyu Zhang , Yixin Wu , Boyang Zhang , Chenhao Lin , Chao Shen , Michael Backes , Yang Zhang

This paper aims to investigate representation learning for large scale visual place recognition, which consists of determining the location depicted in a query image by referring to a database of reference images. This is a challenging task…

Computer Vision and Pattern Recognition · Computer Science 2022-10-20 Amar Ali-bey , Brahim Chaib-draa , Philippe Giguère

The image geolocalization task aims to predict the location where an image was taken anywhere on Earth using visual clues. Existing large vision-language model (LVLM) approaches leverage world knowledge, chain-of-thought reasoning, and…

Computer Vision and Pattern Recognition · Computer Science 2026-01-12 Yuxiang Ji , Yong Wang , Ziyu Ma , Yiming Hu , Hailang Huang , Xuecai Hu , Guanhua Chen , Liaoni Wu , Xiangxiang Chu

We propose a pipeline for combined multi-class object geolocation and height estimation from street level RGB imagery, which is considered as a single available input data modality. Our solution is formulated via Markov Random Field…

Computer Vision and Pattern Recognition · Computer Science 2023-05-16 Matej Ulicny , Vladimir A. Krylov , Julie Connelly , Rozenn Dahyot

We present a novel generative method for the creation of city-scale road layouts. While the output of recent methods is limited in both size of the covered area and diversity, our framework produces large traversable graphs of high quality…

Machine Learning · Computer Science 2022-09-02 Michael Birsak , Tom Kelly , Wamiq Para , Peter Wonka

Planet-scale photo geolocalization is the complex task of estimating the location depicted in an image solely based on its visual content. Due to the success of convolutional neural networks (CNNs), current approaches achieve super-human…

Computer Vision and Pattern Recognition · Computer Science 2021-10-22 Jonas Theiner , Eric Mueller-Budack , Ralph Ewerth

This paper explores the concept of leveraging generative AI as a mapping assistant for enhancing the efficiency of collaborative mapping. We present results of an experiment that combines multiple sources of volunteered geographic…

Computers and Society · Computer Science 2024-03-18 Levente Juhász , Peter Mooney , Hartwig H. Hochmair , Boyuan Guan

Multi-modal large language models have demonstrated impressive performance across various tasks in different modalities. However, existing multi-modal models primarily emphasize capturing global information within each modality while…

Computer Vision and Pattern Recognition · Computer Science 2024-03-06 Zhaowei Li , Qi Xu , Dong Zhang , Hang Song , Yiqing Cai , Qi Qi , Ran Zhou , Junting Pan , Zefeng Li , Van Tu Vu , Zhida Huang , Tao Wang

Multimodal large language models (MLLMs), such as GPT-4o, Gemini, LLaVA, and Flamingo, have made significant progress in integrating visual and textual modalities, excelling in tasks like visual question answering (VQA), image captioning,…

Computer Vision and Pattern Recognition · Computer Science 2024-12-31 Junxiao Xue , Quan Deng , Fei Yu , Yanhao Wang , Jun Wang , Yuehua Li

This paper describes a multi-modal data association method for global localization using object-based maps and camera images. In global localization, or relocalization, using object-based maps, existing methods typically resort to matching…

Computer Vision and Pattern Recognition · Computer Science 2024-02-12 Shigemichi Matsuzaki , Takuma Sugino , Kazuhito Tanaka , Zijun Sha , Shintaro Nakaoka , Shintaro Yoshizawa , Kazuhiro Shintani

The advancement of Multimodal Large Language Models (MLLMs) has greatly accelerated the development of applications in understanding integrated texts and images. Recent works leverage image-caption datasets to train MLLMs, achieving…

Computation and Language · Computer Science 2024-11-22 Mingxu Tao , Quzhe Huang , Kun Xu , Liwei Chen , Yansong Feng , Dongyan Zhao

We propose to use deep convolutional neural networks to address the problem of cross-view image geolocalization, in which the geolocation of a ground-level query image is estimated by matching to georeferenced aerial images. We use…

Computer Vision and Pattern Recognition · Computer Science 2015-10-14 Scott Workman , Richard Souvenir , Nathan Jacobs

Millions of biological sample records collected in the last few centuries archived in natural history collections are un-georeferenced. Georeferencing complex locality descriptions associated with these collection samples is a highly…

Artificial Intelligence · Computer Science 2025-07-14 Kalana Wijegunarathna , Kristin Stock , Christopher B. Jones

The choice of representation for geographic location significantly impacts the accuracy of models for a broad range of geospatial tasks, including fine-grained species classification, population density estimation, and biome classification.…

Computer Vision and Pattern Recognition · Computer Science 2025-04-07 Aayush Dhakal , Srikumar Sastry , Subash Khanal , Adeel Ahmad , Eric Xing , Nathan Jacobs

The ability to recognize the position and order of the floor-level lines that divide adjacent building floors can benefit many applications, for example, urban augmented reality (AR). This work tackles the problem of locating floor-level…

Computer Vision and Pattern Recognition · Computer Science 2021-08-11 Mengyang Wu , Wei Zeng , Chi-Wing Fu

Fine-grained recognition distinguishes among categories with subtle visual differences. In order to differentiate between these challenging visual categories, it is helpful to leverage additional information. Geolocation is a rich source of…

Computer Vision and Pattern Recognition · Computer Science 2019-09-06 Grace Chu , Brian Potetz , Weijun Wang , Andrew Howard , Yang Song , Fernando Brucher , Thomas Leung , Hartwig Adam

As Large Language Models (LLMs) become popular, there emerged an important trend of using multimodality to augment the LLMs' generation ability, which enables LLMs to better interact with the world. However, there lacks a unified perception…

Computation and Language · Computer Science 2023-12-04 Ruochen Zhao , Hailin Chen , Weishi Wang , Fangkai Jiao , Xuan Long Do , Chengwei Qin , Bosheng Ding , Xiaobao Guo , Minzhi Li , Xingxuan Li , Shafiq Joty

Image geolocalization is the challenging task of predicting the geographic coordinates of origin for a given photo. It is an unsolved problem relying on the ability to combine visual clues with general knowledge about the world to make…

Computer Vision and Pattern Recognition · Computer Science 2023-02-02 Lukas Haas , Silas Alberti , Michal Skreta