English
Related papers

Related papers: TerraMind: Large-Scale Generative Multimodality fo…

200 papers

Earth observation is becoming one of the largest data-producing activities in science, yet current pipelines still treat compression as a storage and transmission tool rather than a new way to use data. We present a generative compression…

Distributed, Parallel, and Cluster Computing · Computer Science 2026-05-12 Jinxiao Zhang , Runmin Dong , Xiyong Wu , Xihan Huang , Shenggan Cheng , Yunkai Yang , Zheng Zhou , Yunpu Xu , Zhaoyang Luo , Miao Yang , Fan Wei , Mengxuan Chen , Yang You , Juepeng Zheng , Weijia Li , Yutong Lu , Haohuan Fu

Humans understand the world through the integration of multiple sensory modalities, enabling them to perceive, reason about, and imagine dynamic physical processes. Inspired by this capability, multimodal foundation models (MFMs) have…

Artificial Intelligence · Computer Science 2025-10-07 Xuehai He

Earth Observation Foundation Models (EOFMs) have exploded in prevalence as tools for processing the massive volumes of remotely sensed and other earth observation data, and for delivering impact on the many essential earth monitoring tasks.…

Computer Vision and Pattern Recognition · Computer Science 2025-10-07 Ryan P. Demilt , Nicholas LaHaye , Karis Tenneson

This paper presents EarthView, a comprehensive dataset specifically designed for self-supervision on remote sensing data, intended to enhance deep learning applications on Earth monitoring tasks. The dataset spans 15 tera pixels of global…

Computer Vision and Pattern Recognition · Computer Science 2025-01-15 Diego Velazquez , Pau Rodriguez López , Sergio Alonso , Josep M. Gonfaus , Jordi Gonzalez , Gerardo Richarte , Javier Marin , Yoshua Bengio , Alexandre Lacoste

The innovative application of precise geospatial vegetation forecasting holds immense potential across diverse sectors, including agriculture, forestry, humanitarian aid, and carbon accounting. To leverage the vast availability of satellite…

Computer Vision and Pattern Recognition · Computer Science 2024-03-08 Vitus Benson , Claire Robin , Christian Requena-Mesa , Lazaro Alonso , Nuno Carvalhais , José Cortés , Zhihan Gao , Nora Linscheid , Mélanie Weynants , Markus Reichstein

Many learning tasks involve multi-modal data streams, where continuous data from different modes convey a comprehensive description about objects. A major challenge in this context is how to efficiently interpret multi-modal information in…

Machine Learning · Computer Science 2020-07-24 Amila Silva , Shanika Karunasekera , Christopher Leckie , Ling Luo

Earth observation (EO) foundation models (FMs) are increasingly trained on multisensor data, spanning multispectral imagery (MSI), synthetic aperture radar (SAR), and derived geospatial layers, but hyperspectral imagery (HSI) remains…

Computer Vision and Pattern Recognition · Computer Science 2026-05-21 Nassim Ait Ali Braham , Aaron Banze , Conrad M. Albrecht , Julien Mairal , Jocelyn Chanussot , Xiao Xiang Zhu

The recent advancement of generative foundational models has ushered in a new era of image generation in the realm of natural images, revolutionizing art design, entertainment, environment simulation, and beyond. Despite producing…

Computer Vision and Pattern Recognition · Computer Science 2024-10-16 Zhiping Yu , Chenyang Liu , Liqin Liu , Zhenwei Shi , Zhengxia Zou

The availability of temporal geospatial data in multiple modalities has been extensively leveraged to enhance the performance of machine learning models. While efforts on the design of adequate model architectures are approaching a level of…

Machine Learning · Computer Science 2024-08-22 Hiba Najjar , Marlon Nuske , Andreas Dengel

The use of robotics in humanitarian demining increasingly involves computer vision techniques to improve landmine detection capabilities. However, in the absence of diverse and realistic datasets, the reliable validation of algorithms…

Existing deep learning methods for remote sensing image fusion often suffer from poor generalization when applied to unseen datasets due to the limited availability of real training data and the domain gap between different satellite…

Computer Vision and Pattern Recognition · Computer Science 2025-12-03 Yongchuan Cui , Peng Liu , Yi Zeng

Large Multimodal Models (LMMs) encode rich factual knowledge via cross-modal pre-training, yet their static representations struggle to maintain an accurate understanding of time-sensitive factual knowledge. Existing benchmarks remain…

Computation and Language · Computer Science 2026-04-08 Kailin Jiang , Ning Jiang , Yuntao Du , Yuchen Ren , Yuchen Li , Yifan Gao , Jinhe Bi , Yunpu Ma , Bin Li , Lei Liu , Qing Li

Earth Observation (EO) analysis is inherently interactive: resolving uncertainty often requires expanding the region of interest, retrieving historical observations, and switching across sensors such as optical and Synthetic Aperture Radar.…

Artificial Intelligence · Computer Science 2026-05-05 Sai Ma , Zhuang Li , Sichao Li , Xinyue Xu , Ruibiao Zhu , Tony Boston , John A. Taylor

In this work, we present a novel method for extensive multi-scale generative terrain modeling. At the core of our model is a cascade of superresolution diffusion models that can be combined to produce consistent images across multiple…

Computer Vision and Pattern Recognition · Computer Science 2024-09-12 Ansh Sharma , Albert Xiao , Praneet Rathi , Rohit Kundu , Albert Zhai , Yuan Shen , Shenlong Wang

In autonomous driving, transparency in the decision-making of perception models is critical, as even a single misperception can be catastrophic. Yet with multi-sensor inputs, it is difficult to determine how each modality contributes to a…

Computer Vision and Pattern Recognition · Computer Science 2025-11-04 Jaehyun Park , Konyul Park , Daehun Kim , Junseo Park , Jun Won Choi

Foundation models have the potential to transform the landscape of remote sensing (RS) data analysis by enabling large computer vision models to be pre-trained on vast amounts of remote sensing data. These models can then be fine-tuned with…

Computer Vision and Pattern Recognition · Computer Science 2024-11-27 Caleb S. Spradlin , Jordan A. Caraballo-Vega , Jian Li , Mark L. Carroll , Jie Gong , Paul M. Montesano

The growing availability of high-quality Earth Observation (EO) data enables accurate global land cover and crop type monitoring. However, the volume and heterogeneity of these datasets pose major processing and annotation challenges. To…

Computer Vision and Pattern Recognition · Computer Science 2026-03-06 Anatol Garioud , Sébastien Giordano , Nicolas David , Nicolas Gonthier

Multimodal large-scale datasets for outdoor scenes are mostly designed for urban driving problems. The scenes are highly structured and semantically different from scenarios seen in nature-centered scenes such as gardens or parks. To…

Computer Vision and Pattern Recognition · Computer Science 2020-11-12 Hoang-An Le , Thomas Mensink , Partha Das , Sezer Karaoglu , Theo Gevers

We present ImageBind, an approach to learn a joint embedding across six different modalities - images, text, audio, depth, thermal, and IMU data. We show that all combinations of paired data are not necessary to train such a joint…

Computer Vision and Pattern Recognition · Computer Science 2023-06-01 Rohit Girdhar , Alaaeldin El-Nouby , Zhuang Liu , Mannat Singh , Kalyan Vasudev Alwala , Armand Joulin , Ishan Misra

This work presents SSL4EO-S12 v1.1, a multimodal, multitemporal Earth Observation dataset designed for pretraining large-scale foundation models. Building on the success of SSL4EO-S12, this extension updates the previous version to fix…

Computer Vision and Pattern Recognition · Computer Science 2026-02-18 Benedikt Blumenstiel , Nassim Ait Ali Braham , Conrad M Albrecht , Stefano Maurogiovanni , Paolo Fraccaro
‹ Prev 1 3 4 5 6 7 10 Next ›