English

UNIGEOCLIP: Unified Geospatial Contrastive Learning

Computer Vision and Pattern Recognition 2026-04-14 v1

Abstract

The growing availability of co-located geospatial data spanning aerial imagery, street-level views, elevation models, text, and geographic coordinates offers a unique opportunity for multimodal representation learning. We introduce UNIGEOCLIP, a massively multimodal contrastive framework to jointly align five complementary geospatial modalities in a single unified embedding space. Unlike prior approaches that fuse modalities or rely on a central pivot representation, our method performs all-to-all contrastive alignment, enabling seamless comparison, retrieval, and reasoning across arbitrary combinations of modalities. We further propose a scaled latitude-longitude encoder that improves spatial representation by capturing multi-scale geographic structure. Extensive experiments across downstream geospatial tasks demonstrate that UNIGEOCLIP consistently outperforms single-modality contrastive models and coordinate-only baselines, highlighting the benefits of holistic multimodal geospatial alignment. A reference implementation is available at https://gastruc.github.io/unigeoclip.

Keywords

Cite

@article{arxiv.2604.11668,
  title  = {UNIGEOCLIP: Unified Geospatial Contrastive Learning},
  author = {Guillaume Astruc and Eduard Trulls and Jan Hosang and Loic Landrieu and Paul-Edouard Sarlin},
  journal= {arXiv preprint arXiv:2604.11668},
  year   = {2026}
}
R2 v1 2026-07-01T12:06:49.736Z