English
Related papers

Related papers: TerraMind: Large-Scale Generative Multimodality fo…

200 papers

Current multimodal and multitask foundation models like 4M or UnifiedIO show promising results, but in practice their out-of-the-box abilities to accept diverse inputs and perform diverse tasks are limited by the (usually rather small)…

Computer Vision and Pattern Recognition · Computer Science 2024-06-17 Roman Bachmann , Oğuzhan Fatih Kar , David Mizrahi , Ali Garjani , Mingfei Gao , David Griffiths , Jiaming Hu , Afshin Dehghan , Amir Zamir

Flood hazard mapping is essential for disaster prevention but remains challenging in data-scarce regions, where traditional hydrodynamic models require extensive geophysical inputs. This paper introduces \textit{ZeroFlood}, a framework that…

Machine Learning · Computer Science 2026-04-01 Hyeongkyun Kim , Orestis Oikonomou

Earth observation (EO) in open-world settings presents a unique challenge: different applications rely on diverse sensor modalities, each with varying ground sampling distances, spectral ranges, and numbers of spectral bands. However,…

Computer Vision and Pattern Recognition · Computer Science 2025-10-17 Zhitong Xiong , Yi Wang , Fahong Zhang , Adam J. Stewart , Joëlle Hanna , Damian Borth , Ioannis Papoutsis , Bertrand Le Saux , Gustau Camps-Valls , Xiao Xiang Zhu

The rapid evolution of multimodal foundation model has demonstrated significant progresses in vision-language understanding and generation, e.g., our previous work SEED-LLaMA. However, there remains a gap between its capability and the…

Computer Vision and Pattern Recognition · Computer Science 2025-03-04 Yuying Ge , Sijie Zhao , Jinguo Zhu , Yixiao Ge , Kun Yi , Lin Song , Chen Li , Xiaohan Ding , Ying Shan

Despite recent progress in Multi-Modal Large Language Models (MLLMs), it remains challenging to integrate diverse tasks ranging from pixel-level perception to high-fidelity generation. Existing approaches often suffer from either restricted…

Computation and Language · Computer Science 2026-01-29 Bin Zhu , Munan Ning , Peng Jin , Bin Lin , Jinfa Huang , Qi Song , Junwu Zhang , Zhenyu Tang , Mingjun Pan , Li Yuan

In this research work, we have proposed a thermal tiny-YOLO multi-class object detection (TTYMOD) system as a smart forward sensing system that should remain effective in all weather and harsh environmental conditions using an end-to-end…

Computer Vision and Pattern Recognition · Computer Science 2023-01-19 Muhammad Ali Farooq , Waseem Shariff , Faisal Khan , Peter Corcoran

This paper presents Prithvi-EO-2.0, a new geospatial foundation model that offers significant improvements over its predecessor, Prithvi-EO-1.0. Trained on 4.2 million global time series samples from NASA's Harmonized Landsat and Sentinel-2…

While integrating multiple modalities has the potential to improve environmental monitoring, current approaches struggle to combine data sources with heterogeneous formats or contents. A central difficulty arises when combining continuous…

Computation and Language · Computer Science 2026-03-27 Valerie Zermatten , Chiara Vanalli , Gencer Sumbul , Diego Marcos , Devis Tuia

Multimodal intent understanding is a significant research area that requires effective leveraging of multiple modalities to analyze human language. Existing methods face two main challenges in this domain. Firstly, they have limitations in…

Multimedia · Computer Science 2025-05-26 Hanlei Zhang , Qianrui Zhou , Hua Xu , Jianhua Su , Roberto Evans , Kai Gao

Deep understanding of electromagnetic signals is fundamental to dynamic spectrum management, intelligent transportation, autonomous driving and unmanned vehicle perception. The field faces challenges because electromagnetic signals differ…

Signal Processing · Electrical Eng. & Systems 2025-08-27 Luqing Luo , Wenjin Gui , Yunfei Liu , Ziyue Zhang , Yunxi Zhang , Fengxiang Wang , Zonghao Guo , Zizhi Ma , Xinzhu Liu , Hanxiang He , Jinhai Li , Xin Qiu , Wupeng Xie , Yangang Sun

Multimodal Large Language Models have demonstrated powerful cross-modal understanding and reasoning capabilities in general domains. However, in the electromagnetic (EM) domain, they still face challenges such as data scarcity and…

Segment Anything Model (SAM) has demonstrated impressive zero-shot segmentation capabilities across natural image domains, but it struggles to generalize to the unique challenges of remote sensing data, such as complex terrain, multi-scale…

Computer Vision and Pattern Recognition · Computer Science 2025-09-22 Tianyang Wang , Xi Xiao , Gaofei Chen , Hanzhang Chi , Qi Zhang , Guo Cheng , Yingrui Ji

We propose MToMnet - a Theory of Mind (ToM) neural network for predicting beliefs and their dynamics during human social interactions from multimodal input. ToM is key for effective nonverbal human communication and collaboration, yet,…

Artificial Intelligence · Computer Science 2024-08-29 Matteo Bortoletto , Constantin Ruhdorfer , Lei Shi , Andreas Bulling

In this paper, we introduce MIO, a novel foundation model built on multimodal tokens, capable of understanding and generating speech, text, images, and videos in an end-to-end, autoregressive manner. While the emergence of large language…

LiDAR perception is fundamental to robotics, enabling machines to understand their environment in 3D. A crucial task for LiDAR-based scene understanding and navigation is ground segmentation. However, existing methods are either handcrafted…

Computer Vision and Pattern Recognition · Computer Science 2026-03-31 Ted Lentsch , Santiago Montiel-Marín , Holger Caesar , Dariu M. Gavrila

Satellite image time series (SITS) provide continuous observations of the Earth's surface, making them essential for applications such as environmental management and disaster assessment. However, existing spatiotemporal foundation models…

Computer Vision and Pattern Recognition · Computer Science 2025-05-14 Xiaolei Qin , Di Wang , Jing Zhang , Fengxiang Wang , Xin Su , Bo Du , Liangpei Zhang

A significant amount of remotely sensed data is generated daily by many Earth observation (EO) spaceborne and airborne sensors over different countries of our planet. Different applications use those data, such as natural hazard monitoring,…

Image and Video Processing · Electrical Eng. & Systems 2024-10-23 Alessandro Sebastianelli , Francesco Mauro , Giulia Ciabatti , Dario Spiller , Bertrand Le Saux , Paolo Gamba , Silvia Ullo

Masked Autoencoders (MAE) play a pivotal role in learning potent representations, delivering outstanding results across various 3D perception tasks essential for autonomous driving. In real-world driving scenarios, it's commonplace to…

Computer Vision and Pattern Recognition · Computer Science 2024-08-26 Jian Zou , Tianyu Huang , Guanglei Yang , Zhenhua Guo , Tao Luo , Chun-Mei Feng , Wangmeng Zuo

Achieving robust spatial reasoning remains a fundamental challenge for current Multimodal Foundation Models (MFMs). Existing methods either overfit statistical shortcuts via 3D grounding data or remain confined to 2D visual perception,…

Artificial Intelligence · Computer Science 2026-03-11 Shouwei Ruan , Bin Wang , Zhenyu Wu , Qihui Zhu , Yuxiang Zhang , Hang Su , Yubin Wang

We present TaxaBind, a unified embedding space for characterizing any species of interest. TaxaBind is a multimodal embedding space across six modalities: ground-level images of species, geographic location, satellite image, text, audio,…

Computer Vision and Pattern Recognition · Computer Science 2024-11-04 Srikumar Sastry , Subash Khanal , Aayush Dhakal , Adeel Ahmad , Nathan Jacobs