English
Related papers

Related papers: TerraMind: Large-Scale Generative Multimodality fo…

200 papers

Multimodal spatiotemporal learning on real-world experimental data is constrained by two challenges: within-modality measurements are sparse, irregular, and noisy (QA/QC artifacts) but cross-modally correlated; the set of available…

Machine Learning · Computer Science 2025-11-05 Kevin Valencia , Thilina Balasooriya , Xihaier Luo , Shinjae Yoo , David Keetae Park

Multimodal models have demonstrated powerful capabilities in complex tasks requiring multimodal alignment, including zero-shot classification and cross-modal retrieval. However, existing models typically rely on millions of paired…

Computer Vision and Pattern Recognition · Computer Science 2025-10-23 Fabian Gröger , Shuo Wen , Huyen Le , Maria Brbić

We present GeoGrid-Bench, a benchmark designed to evaluate the ability of foundation models to understand geo-spatial data in the grid structure. Geo-spatial datasets pose distinct challenges due to their dense numerical values, strong…

Computation and Language · Computer Science 2025-05-27 Bowen Jiang , Yangxinyu Xie , Xiaomeng Wang , Jiashu He , Joshua Bergerson , John K Hutchison , Jordan Branham , Camillo J Taylor , Tanwi Mallick

To make sense of their surroundings, intelligent systems must transform complex sensory inputs to structured codes that are reduced to task-relevant information such as object category. Biological agents achieve this in a largely autonomous…

Computer Vision and Pattern Recognition · Computer Science 2024-07-09 Robin Weiler , Matthias Brucklacher , Cyriel M. A. Pennartz , Sander M. Bohté

The rapid evolution of satellite-borne Earth Observation (EO) systems has revolutionized terrestrial monitoring, yielding petabyte-scale archives. However, the immense computational and storage requirements for global-scale analysis often…

Computer Vision and Pattern Recognition · Computer Science 2026-01-19 Shuang Chen , Jie Wang , Shuai Yuan , Jiayang Li , Yu Xia , Yuanhong Liao , Junbo Wei , Jincheng Yuan , Xiaoqing Xu , Xiaolin Zhu , Peng Zhu , Hongsheng Zhang , Yuyu Zhou , Haohuan Fu , Huabing Huang , Bin Chen , Fan Dai , Peng Gong

A robust Multimodal Large Language Model (MLLM) for Earth Observation should maintain consistent interpretation and reasoning under realistic input variations. However, current Remote Sensing MLLMs fail to meet this requirement. Trained on…

Computer Vision and Pattern Recognition · Computer Science 2026-04-21 Rui Min , Liang Yao , Shiyu Miao , Shengxiang Xu , Yuxuan Liu , Chuanyi Zhang , Shimin Di , Fan Liu

Geospatial foundation models (GeoFMs) promise broad generalisation capacity for Earth observation (EO) tasks, particularly under data-limited conditions. However, their large size poses a barrier to deployment on resource-constrained space…

Multimodal multitask learning has attracted an increasing interest in recent years. Singlemodal models have been advancing rapidly and have achieved astonishing results on various tasks across multiple domains. Multimodal learning offers…

Computer Vision and Pattern Recognition · Computer Science 2023-07-04 Ye Xue , Diego Klabjan , Jean Utke

Existing benchmarks for multimodal learning in Earth science offer limited, siloed coverage of Earth's spheres and their cross-sphere interactions, typically restricting evaluation to the human-activity sphere of atmosphere and to at most…

Computer Vision and Pattern Recognition · Computer Science 2026-02-17 Fengxiang Wang , Mingshuo Chen , Xuming He , Yi-Fan Zhang , Yueying Li , Feng Liu , Zijie Guo , Zhenghao Hu , Jiong Wang , Jingyi Xu , Zhangrui Li , Junchao Gong , Di Wang , Fenghua Ling , Ben Fei , Weijia Li , Long Lan , Wenjing Yang

Recent progress in unified models for image understanding and generation has been impressive, yet most approaches remain limited to single-modal generation conditioned on multiple modalities. In this paper, we present Mogao, a unified…

Computer Vision and Pattern Recognition · Computer Science 2025-05-13 Chao Liao , Liyang Liu , Xun Wang , Zhengxiong Luo , Xinyu Zhang , Wenliang Zhao , Jie Wu , Liang Li , Zhi Tian , Weilin Huang

Emotion recognition based on Electroencephalography (EEG) has gained significant attention and diversified development in fields such as neural signal processing and affective computing. However, the unique brain anatomy of individuals…

Signal Processing · Electrical Eng. & Systems 2024-05-31 Yihang Dong , Xuhang Chen , Yanyan Shen , Michael Kwok-Po Ng , Tao Qian , Shuqiang Wang

Despite the unprecedented volume of multimodal data provided by modern Earth observation systems, our ability to model atmospheric dynamics remains constrained. Traditional modeling frameworks force heterogeneous measurements into…

Recent deep learning models can efficiently combine inputs from different modalities (e.g., images and text) and learn to align their latent representations, or to translate signals from one domain to another (as in image captioning, or…

Artificial Intelligence · Computer Science 2025-11-27 Benjamin Devillers , Léopold Maytié , Rufin VanRullen

Building a multi-modality multi-task neural network toward accurate and robust performance is a de-facto standard in perception task of autonomous driving. However, leveraging such data from multiple sensors to jointly optimize the…

Computer Vision and Pattern Recognition · Computer Science 2023-08-15 Tengju Ye , Wei Jing , Chunyong Hu , Shikun Huang , Lingping Gao , Fangzhen Li , Jingke Wang , Ke Guo , Wencong Xiao , Weibo Mao , Hang Zheng , Kun Li , Junbo Chen , Kaicheng Yu

Multimodal interleaved datasets featuring free-form interleaved sequences of images and text are crucial for training frontier large multimodal models (LMMs). Despite the rapid progression of open-source LMMs, there remains a pronounced…

Computer Vision and Pattern Recognition · Computer Science 2024-11-01 Anas Awadalla , Le Xue , Oscar Lo , Manli Shu , Hannah Lee , Etash Kumar Guha , Matt Jordan , Sheng Shen , Mohamed Awadalla , Silvio Savarese , Caiming Xiong , Ran Xu , Yejin Choi , Ludwig Schmidt

Understanding and predicting urban dynamics is crucial for managing transportation systems, optimizing urban planning, and enhancing public services. While neural network-based approaches have achieved success, they often rely on…

Machine Learning · Computer Science 2025-05-23 Yuhang Liu , Yingxue Zhang , Xin Zhang , Ling Tian , Yanhua Li , Jun Luo

Multi-view learning (MVL) leverages multiple sources or views of data to enhance machine learning model performance and robustness. This approach has been successfully used in the Earth Observation (EO) domain, where views have a…

Machine Learning · Computer Science 2025-09-12 Francisco Mena , Diego Arenas , Andreas Dengel

Efficient observer design and accurate sensor fusion are key in state estimation. This work proposes an optimization-based methodology, termed Trajectory Based Optimization Design (TBOD), allowing the user to easily design observers for…

Robotics · Computer Science 2025-10-02 Federico Oliva , Tom Shaked , Daniele Carnevale , Amir Degani

While existing time series foundation models primarily rely on large-scale unimodal pretraining, they lack complementary modalities to enhance time series understanding. Building multimodal foundation models is a natural next step, but it…

Machine Learning · Computer Science 2026-02-06 Peng Chen , Siyuan Wang , Shiyan Hu , Xingjian Wu , Yang Shu , Zhongwen Rao , Meng Wang , Yijie Li , Bin Yang , Chenjuan Guo

Current large multimodal models (LMMs) face challenges in grounding, which requires the model to relate language components to visual entities. Contrary to the common practice that fine-tunes LMMs with additional grounding supervision, we…

Computer Vision and Pattern Recognition · Computer Science 2025-10-17 Shengcao Cao , Liang-Yan Gui , Yu-Xiong Wang