Image geolocation is a critical task in various image-understanding applications. However, existing methods often fail when analyzing challenging, in-the-wild images. Inspired by the exceptional background knowledge of multimodal language models, we systematically evaluate their geolocation capabilities using a novel image dataset and a comprehensive evaluation framework. We first collect images from various countries via Google Street View. Then, we conduct training-free and training-based evaluations on closed-source and open-source multi-modal language models. we conduct both training-free and training-based evaluations on closed-source and open-source multimodal language models. Our findings indicate that closed-source models demonstrate superior geolocation abilities, while open-source models can achieve comparable performance through fine-tuning.
@article{arxiv.2405.20363,
title = {LLMGeo: Benchmarking Large Language Models on Image Geolocation In-the-wild},
author = {Zhiqiang Wang and Dejia Xu and Rana Muhammad Shahroz Khan and Yanbin Lin and Zhiwen Fan and Xingquan Zhu},
journal= {arXiv preprint arXiv:2405.20363},
year = {2024}
}
Comments
7 pages, 3 figures, 5 tables, CVPR 2024 Workshop on Computer Vision in the Wild