中文
相关论文

相关论文: OlmoEarth: Stable Latent Image Modeling for Multim…

200 篇论文

Foundation models have transformed natural language processing and computer vision, and their impact is now reshaping remote sensing image analysis. With powerful generalization and transfer learning capabilities, they align naturally with…

计算机视觉与模式识别 · 计算机科学 2025-10-28 Liling Yang , Ning Chen , Jun Yue , Yidan Liu , Jiayi Ma , Pedram Ghamisi , Antonio Plaza , Leyuan Fang

Machine learning models deployed in real-world settings must operate under evolving data distributions and constrained computational resources. This challenge is particularly acute in non-stationary domains such as energy time series,…

机器学习 · 计算机科学 2026-03-17 Daniel Bretsko , Piotr Walas , Devashish Khulbe , Sebastian Stros , Stanislav Sobolevsky , Tomas Satura

Foundation models that incorporate language, vision, and more recently actions have revolutionized the ability to harness internet scale data to reason about useful tasks. However, one of the key challenges of training embodied foundation…

Tracking a time-varying indefinite number of objects in a video sequence over time remains a challenge despite recent advances in the field. Most existing approaches are not able to properly handle multi-object tracking challenges such as…

计算机视觉与模式识别 · 计算机科学 2022-10-10 Tianyu Zhu , Markus Hiller , Mahsa Ehsanpour , Rongkai Ma , Tom Drummond , Ian Reid , Hamid Rezatofighi

Carefully curated and annotated datasets are the foundation of machine learning, with particularly data-hungry deep neural networks forming the core of what is often called Artificial Intelligence (AI). Due to the massive success of deep…

计算机视觉与模式识别 · 计算机科学 2023-10-31 Michael Schmitt , Seyed Ali Ahmadi , Yonghao Xu , Gulsen Taskin , Ujjwal Verma , Francescopaolo Sica , Ronny Hansch

Satellite-based remote sensing has revolutionised the way we address global challenges. Huge quantities of Earth Observation (EO) data are generated by satellite sensors daily, but processing these large datasets for use in ML pipelines is…

We develop a foundation model using 1.2m high resolution satellite images of the Netherlands. By combining a Convolutional Neural Network and a Vision Transformer, the model captures both low- and high-frequency landscape features, such as…

计算机视觉与模式识别 · 计算机科学 2026-05-12 Paul Vermeeren , Heysem Kaya

Geo-spatial analysis of our world benefits from a multimodal approach, as every single geographic location can be described in numerous ways (images from various viewpoints, textual descriptions, geographic coordinates, etc.). Current…

计算机视觉与模式识别 · 计算机科学 2026-04-29 Oskar Kristoffersen , Alba Reinders Sánchez , Morten Rieger Hannemose , Anders Bjorholm Dahl , Dim P. Papadopoulos

We propose a task-agnostic framework for multimodal fusion of time series and single timestamp images, enabling cross-modal generation and robust downstream performance. Our approach explores deterministic and learned strategies for time…

计算机视觉与模式识别 · 计算机科学 2026-01-27 Gianfranco Basile , Johannes Jakubik , Benedikt Blumenstiel , Thomas Brunschwiler , Juan Bernabe Moreno

In recent years, the development of robust multi-source models has emerged in the Earth Observation (EO) field. These are models that leverage data from diverse sources to improve predictive accuracy when there is missing data. Despite…

机器学习 · 计算机科学 2026-05-14 Francisco Mena , Diego Arenas , Miro Miranda , Andreas Dengel

Ultra-high-resolution (UHR) remote sensing (RS) images offer rich fine-grained information but also present challenges in effective processing. Existing dynamic resolution and token pruning methods are constrained by a passive perception…

计算机视觉与模式识别 · 计算机科学 2025-11-18 Ruixun Liu , Bowen Fu , Jiayi Song , Kaiyu Li , Wanchen Li , Lanxuan Xue , Hui Qiao , Weizhan Zhang , Deyu Meng , Xiangyong Cao

We introduce MOMO, the first multi-sensor foundation model for Mars remote sensing. MOMO uses model merge to integrate representations learned independently from three key Martian sensors (HiRISE, CTX, and THEMIS), spanning resolutions from…

计算机视觉与模式识别 · 计算机科学 2026-04-06 Mirali Purohit , Bimal Gajera , Irish Mehta , Bhanu Tokas , Jacob Adler , Steven Lu , Scott Dickenshied , Serina Diniega , Brian Bue , Umaa Rebbapragada , Hannah Kerner

Geospatial foundation models (GFMs) for Earth observation often fail to perform reliably in environments underrepresented during pretraining. We introduce SHRUG-FM, a framework for reliability-aware prediction that enables GFMs to identify…

We introduce EarthPT -- an Earth Observation (EO) pretrained transformer. EarthPT is a 700 million parameter decoding transformer foundation model trained in an autoregressive self-supervised manner and developed specifically with EO…

机器学习 · 计算机科学 2024-01-12 Michael J. Smith , Luke Fleming , James E. Geach

Pretrained models have demonstrated impressive success in many modalities such as language and vision. Recent works facilitate the pretraining paradigm in imaging research. Transients are a novel modality, which are captured for an object…

计算机视觉与模式识别 · 计算机科学 2025-06-11 Siyuan Shen , Ziheng Wang , Xingyue Peng , Suan Xia , Ruiqian Li , Shiying Li , Jingyi Yu

Object detection in remote sensing imagery plays a vital role in various Earth observation applications. However, unlike object detection in natural scene images, this task is particularly challenging due to the abundance of small, often…

计算机视觉与模式识别 · 计算机科学 2024-09-16 Minh-Duc Vu , Zuheng Ming , Fangchen Feng , Bissmella Bahaduri , Anissa Mokraoui

Off-road environments remain significant challenges for autonomous ground vehicles, due to the lack of structured roads and the presence of complex obstacles, such as uneven terrain, vegetation, and occlusions. Traditional perception…

机器人学 · 计算机科学 2025-08-07 Zitong Chen , Chao Sun , Shida Nie , Chen Min , Changjiu Ning , Haoyu Li , Bo Wang

Multimodal spatiotemporal learning on real-world experimental data is constrained by two challenges: within-modality measurements are sparse, irregular, and noisy (QA/QC artifacts) but cross-modally correlated; the set of available…

机器学习 · 计算机科学 2025-11-05 Kevin Valencia , Thilina Balasooriya , Xihaier Luo , Shinjae Yoo , David Keetae Park

With the ever-increasing volumes of the Earth observation data present in the archives of large programmes such as Copernicus, there is a growing need for efficient vector representations of the underlying raw data. The approach of…

计算机视觉与模式识别 · 计算机科学 2024-12-10 Mikolaj Czerkawski , Marcin Kluczek , Jędrzej S. Bojanowski

The For\^et Montmorency (FoMo) dataset is a comprehensive multi-season data collection, recorded over the span of one year in a boreal forest. Featuring a unique combination of on- and off-pavement environments with significant…