中文
相关论文

相关论文: GeoViSTA: Geospatial Vision-Tabular Transformer fo…

200 篇论文

Spatial intelligence is the ability of a machine to perceive, reason, and act in three dimensions within space and time. Recent advancements in large-scale auto-regressive models have demonstrated remarkable capabilities across various…

计算机视觉与模式识别 · 计算机科学 2024-10-25 Junyi Chen , Di Huang , Weicai Ye , Wanli Ouyang , Tong He

This paper aims to investigate representation learning for large scale visual place recognition, which consists of determining the location depicted in a query image by referring to a database of reference images. This is a challenging task…

计算机视觉与模式识别 · 计算机科学 2022-10-20 Amar Ali-bey , Brahim Chaib-draa , Philippe Giguère

We introduce a highly multimodal transformer to represent many remote sensing modalities - multispectral optical, synthetic aperture radar, elevation, weather, pseudo-labels, and more - across space and time. These inputs are useful for…

计算机视觉与模式识别 · 计算机科学 2025-06-05 Gabriel Tseng , Anthony Fuller , Marlena Reil , Henry Herzog , Patrick Beukema , Favyen Bastani , James R. Green , Evan Shelhamer , Hannah Kerner , David Rolnick

Building spatiotemporal activity models for people's activities in urban spaces is important for understanding the ever-increasing complexity of urban dynamics. With the emergence of Geo-Tagged Social Media (GTSM) records, previous studies…

机器学习 · 计算机科学 2019-10-24 Amila Silva , Shanika Karunasekera , Christopher Leckie , Ling Luo

World-model-based imagine-then-act becomes a promising paradigm for robotic manipulation, yet existing approaches typically support either purely image-based forecasting or reasoning over partial 3D geometry, limiting their ability to…

计算机视觉与模式识别 · 计算机科学 2026-05-27 Jiaxu Wang , Yicheng Jiang , Tianlun He , Jingkai Sun , Qiang Zhang , Junhao He , Jiahang Cao , Zesen Gan , Mingyuan Sun , Qiming Shao , Xiangyu Yue

Self-supervised pre-training based on next-token prediction has enabled large language models to capture the underlying structure of text, and has led to unprecedented performance on a large array of tasks when applied at scale. Similarly,…

Learning causal structures from observational data remains a fundamental yet computationally intensive task, particularly in high-dimensional settings where existing methods face challenges such as the super-exponential growth of the search…

机器学习 · 统计学 2026-02-12 Haixiang Sun , Pengchao Tian , Zihan Zhou , Jielei Zhang , Peiyi Li , Andrew L. Liu

Robust 3D representation learning forms the perceptual foundation of spatial intelligence, enabling downstream tasks in scene understanding and embodied AI. However, learning such representations directly from unposed multi-view images…

计算机视觉与模式识别 · 计算机科学 2026-04-14 Bo Zhou , Qiuxia Lai , Zeren Sun , Xiangbo Shu , Yazhou Yao , Wenguan Wang

Geospatial Information Systems are used by researchers and Humanitarian Assistance and Disaster Response (HADR) practitioners to support a wide variety of important applications. However, collaboration between these actors is difficult due…

Visual representation learning has been a cornerstone in computer vision, involving typical forms such as visual embeddings, structural symbols, and text-based representations. Despite the success of CLIP-type visual embeddings, they often…

计算机视觉与模式识别 · 计算机科学 2024-06-18 Yiwu Zhong , Zi-Yuan Hu , Michael R. Lyu , Liwei Wang

Forecasting high-resolution land subsidence is a critical yet challenging task due to its complex, non-linear dynamics. While standard architectures like ConvLSTM often fail to model long-range dependencies, we argue that a more fundamental…

计算机视觉与模式识别 · 计算机科学 2025-10-02 Wendong Yao , Binhua Huang , Soumyabrata Dev

Cross-view object geo-localization enables high-precision object localization through cross-view matching, with critical applications in autonomous driving, urban management, and disaster response. However, existing methods rely on…

计算机视觉与模式识别 · 计算机科学 2025-10-24 Shuhan Hu , Yiru Li , Yuanyuan Li , Yingying Zhu

We propose PyViT-FUSE, a foundation model for earth observation data explicitly designed to handle multi-modal imagery by learning to fuse an arbitrary number of mixed-resolution input bands into a single representation through an attention…

计算机视觉与模式识别 · 计算机科学 2025-04-29 Manuel Weber , Carly Beneke

This paper addresses the problem of cross-view image geo-localization, where the geographic location of a ground-level street-view query image is estimated by matching it against a large scale aerial map (e.g., a high-resolution satellite…

计算机视觉与模式识别 · 计算机科学 2019-11-28 Yujiao Shi , Xin Yu , Liu Liu , Tong Zhang , Hongdong Li

Spatio-temporal prediction is a crucial research area in data-driven urban computing, with implications for transportation, public safety, and environmental monitoring. However, scalability and generalization challenges remain significant…

机器学习 · 计算机科学 2024-09-12 Jiabin Tang , Wei Wei , Lianghao Xia , Chao Huang

This paper presents GeoDecoder, a dedicated multimodal model designed for processing geospatial information in maps. Built on the BeitGPT architecture, GeoDecoder incorporates specialized expert modules for image and text processing. On the…

计算机视觉与模式识别 · 计算机科学 2024-02-20 Feng Qi , Mian Dai , Zixian Zheng , Chao Wang

We study the image-based geolocalization problem, aiming to localize ground-view query images on cartographic maps. Current methods often utilize cross-view localization techniques to match ground-view query images with 2D maps. However,…

计算机视觉与模式识别 · 计算机科学 2023-11-06 Mengjie Zhou , Liu Liu , Yiran Zhong , Andrew Calway

Large-scale foundation models in Earth Observation can learn versatile, label-efficient representations by leveraging massive amounts of unlabeled data. However, existing public datasets are often limited in scale, geographic coverage, or…

Geospatial imaging leverages data from diverse sensing modalities-such as EO, SAR, and LiDAR, ranging from ground-level drones to satellite views. These heterogeneous inputs offer significant opportunities for scene understanding but…

计算机视觉与模式识别 · 计算机科学 2025-01-20 Alex Berian , Daniel Brignac , JhihYang Wu , Natnael Daba , Abhijit Mahalanobis

Vision-based approaches have become the dominant paradigm for traversability estimation in unstructured outdoor environments, typically adapting vision foundation models (VFMs) via semantic segmentation supervision. However, this paradigm…

计算机视觉与模式识别 · 计算机科学 2026-05-29 Ji-Hoon Hwang , Jisung Bae , Dong-Wook Kim , Yeonkyu Lee , Seung-Woo Seo