English
Related papers

Related papers: GeoViSTA: Geospatial Vision-Tabular Transformer fo…

200 papers

Spatial intelligence is the ability of a machine to perceive, reason, and act in three dimensions within space and time. Recent advancements in large-scale auto-regressive models have demonstrated remarkable capabilities across various…

Computer Vision and Pattern Recognition · Computer Science 2024-10-25 Junyi Chen , Di Huang , Weicai Ye , Wanli Ouyang , Tong He

This paper aims to investigate representation learning for large scale visual place recognition, which consists of determining the location depicted in a query image by referring to a database of reference images. This is a challenging task…

Computer Vision and Pattern Recognition · Computer Science 2022-10-20 Amar Ali-bey , Brahim Chaib-draa , Philippe Giguère

We introduce a highly multimodal transformer to represent many remote sensing modalities - multispectral optical, synthetic aperture radar, elevation, weather, pseudo-labels, and more - across space and time. These inputs are useful for…

Computer Vision and Pattern Recognition · Computer Science 2025-06-05 Gabriel Tseng , Anthony Fuller , Marlena Reil , Henry Herzog , Patrick Beukema , Favyen Bastani , James R. Green , Evan Shelhamer , Hannah Kerner , David Rolnick

Building spatiotemporal activity models for people's activities in urban spaces is important for understanding the ever-increasing complexity of urban dynamics. With the emergence of Geo-Tagged Social Media (GTSM) records, previous studies…

Machine Learning · Computer Science 2019-10-24 Amila Silva , Shanika Karunasekera , Christopher Leckie , Ling Luo

World-model-based imagine-then-act becomes a promising paradigm for robotic manipulation, yet existing approaches typically support either purely image-based forecasting or reasoning over partial 3D geometry, limiting their ability to…

Computer Vision and Pattern Recognition · Computer Science 2026-05-27 Jiaxu Wang , Yicheng Jiang , Tianlun He , Jingkai Sun , Qiang Zhang , Junhao He , Jiahang Cao , Zesen Gan , Mingyuan Sun , Qiming Shao , Xiangyu Yue

Self-supervised pre-training based on next-token prediction has enabled large language models to capture the underlying structure of text, and has led to unprecedented performance on a large array of tasks when applied at scale. Similarly,…

Learning causal structures from observational data remains a fundamental yet computationally intensive task, particularly in high-dimensional settings where existing methods face challenges such as the super-exponential growth of the search…

Machine Learning · Statistics 2026-02-12 Haixiang Sun , Pengchao Tian , Zihan Zhou , Jielei Zhang , Peiyi Li , Andrew L. Liu

Robust 3D representation learning forms the perceptual foundation of spatial intelligence, enabling downstream tasks in scene understanding and embodied AI. However, learning such representations directly from unposed multi-view images…

Computer Vision and Pattern Recognition · Computer Science 2026-04-14 Bo Zhou , Qiuxia Lai , Zeren Sun , Xiangbo Shu , Yazhou Yao , Wenguan Wang

Geospatial Information Systems are used by researchers and Humanitarian Assistance and Disaster Response (HADR) practitioners to support a wide variety of important applications. However, collaboration between these actors is difficult due…

Visual representation learning has been a cornerstone in computer vision, involving typical forms such as visual embeddings, structural symbols, and text-based representations. Despite the success of CLIP-type visual embeddings, they often…

Computer Vision and Pattern Recognition · Computer Science 2024-06-18 Yiwu Zhong , Zi-Yuan Hu , Michael R. Lyu , Liwei Wang

Forecasting high-resolution land subsidence is a critical yet challenging task due to its complex, non-linear dynamics. While standard architectures like ConvLSTM often fail to model long-range dependencies, we argue that a more fundamental…

Computer Vision and Pattern Recognition · Computer Science 2025-10-02 Wendong Yao , Binhua Huang , Soumyabrata Dev

Cross-view object geo-localization enables high-precision object localization through cross-view matching, with critical applications in autonomous driving, urban management, and disaster response. However, existing methods rely on…

Computer Vision and Pattern Recognition · Computer Science 2025-10-24 Shuhan Hu , Yiru Li , Yuanyuan Li , Yingying Zhu

We propose PyViT-FUSE, a foundation model for earth observation data explicitly designed to handle multi-modal imagery by learning to fuse an arbitrary number of mixed-resolution input bands into a single representation through an attention…

Computer Vision and Pattern Recognition · Computer Science 2025-04-29 Manuel Weber , Carly Beneke

This paper addresses the problem of cross-view image geo-localization, where the geographic location of a ground-level street-view query image is estimated by matching it against a large scale aerial map (e.g., a high-resolution satellite…

Computer Vision and Pattern Recognition · Computer Science 2019-11-28 Yujiao Shi , Xin Yu , Liu Liu , Tong Zhang , Hongdong Li

Spatio-temporal prediction is a crucial research area in data-driven urban computing, with implications for transportation, public safety, and environmental monitoring. However, scalability and generalization challenges remain significant…

Machine Learning · Computer Science 2024-09-12 Jiabin Tang , Wei Wei , Lianghao Xia , Chao Huang

This paper presents GeoDecoder, a dedicated multimodal model designed for processing geospatial information in maps. Built on the BeitGPT architecture, GeoDecoder incorporates specialized expert modules for image and text processing. On the…

Computer Vision and Pattern Recognition · Computer Science 2024-02-20 Feng Qi , Mian Dai , Zixian Zheng , Chao Wang

We study the image-based geolocalization problem, aiming to localize ground-view query images on cartographic maps. Current methods often utilize cross-view localization techniques to match ground-view query images with 2D maps. However,…

Computer Vision and Pattern Recognition · Computer Science 2023-11-06 Mengjie Zhou , Liu Liu , Yiran Zhong , Andrew Calway

Large-scale foundation models in Earth Observation can learn versatile, label-efficient representations by leveraging massive amounts of unlabeled data. However, existing public datasets are often limited in scale, geographic coverage, or…

Geospatial imaging leverages data from diverse sensing modalities-such as EO, SAR, and LiDAR, ranging from ground-level drones to satellite views. These heterogeneous inputs offer significant opportunities for scene understanding but…

Computer Vision and Pattern Recognition · Computer Science 2025-01-20 Alex Berian , Daniel Brignac , JhihYang Wu , Natnael Daba , Abhijit Mahalanobis

Vision-based approaches have become the dominant paradigm for traversability estimation in unstructured outdoor environments, typically adapting vision foundation models (VFMs) via semantic segmentation supervision. However, this paradigm…

Computer Vision and Pattern Recognition · Computer Science 2026-05-29 Ji-Hoon Hwang , Jisung Bae , Dong-Wook Kim , Yeonkyu Lee , Seung-Woo Seo