English
Related papers

Related papers: GeoMeld: Toward Semantically Grounded Foundation M…

200 papers

Referring remote sensing image segmentation is crucial for achieving fine-grained visual understanding through free-format textual input, enabling enhanced scene and object extraction in remote sensing applications. Current research…

Computer Vision and Pattern Recognition · Computer Science 2025-01-14 Keyan Chen , Jiafan Zhang , Chenyang Liu , Zhengxia Zou , Zhenwei Shi

We introduce a highly multimodal transformer to represent many remote sensing modalities - multispectral optical, synthetic aperture radar, elevation, weather, pseudo-labels, and more - across space and time. These inputs are useful for…

Computer Vision and Pattern Recognition · Computer Science 2025-06-05 Gabriel Tseng , Anthony Fuller , Marlena Reil , Henry Herzog , Patrick Beukema , Favyen Bastani , James R. Green , Evan Shelhamer , Hannah Kerner , David Rolnick

Earth observation (EO) foundation models have emerged as an effective approach to derive latent representations of the Earth system from various remote sensing sensors. These models produce embeddings that can be used as analysis-ready…

Machine Learning · Computer Science 2025-11-21 Julia Peters , Karin Mora , Miguel D. Mahecha , Chaonan Ji , David Montero , Clemens Mosig , Guido Kraemer

Combining multimodal data is a key issue in a wide range of machine learning tasks, including many remote sensing problems. In Earth observation, early multimodal data fusion methods were based on specific neural network architectures and…

Computer Vision and Pattern Recognition · Computer Science 2025-09-23 Romain Thoreau , Jessie Levillain , Dawa Derksen

Multimodal Large Language Models (MLLMs) have made significant advancements, demonstrating powerful capabilities in processing and understanding multimodal data. Fine-tuning MLLMs with Federated Learning (FL) allows for expanding the…

Machine Learning · Computer Science 2025-03-11 Binqian Xu , Xiangbo Shu , Haiyang Mei , Guosen Xie , Basura Fernando , Jinhui Tang

Visual transformers have driven major progress in remote sensing image analysis, particularly in object detection and segmentation. Recent vision-language and multimodal models further extend these capabilities by incorporating auxiliary…

Computer Vision and Pattern Recognition · Computer Science 2026-02-03 Yu Li , Guilherme N. DeSouza , Praveen Rao , Chi-Ren Shyu

Semantic segmentation is essential for analyzing highdefinition remote sensing images (HRSIs) because it allows the precise classification of objects and regions at the pixel level. However, remote sensing data present challenges owing to…

Computer Vision and Pattern Recognition · Computer Science 2025-03-12 Sachin Verma , Frank Lindseth , Gabriel Kiss

Multi-modal data abounds in biomedicine, such as radiology images and reports. Interpreting this data at scale is essential for improving clinical care and accelerating clinical research. Biomedical text with its complex semantics poses…

Remote sensing (RS) techniques are increasingly crucial for deepening our understanding of the planet. As the volume and diversity of RS data continue to grow exponentially, there is an urgent need for advanced data modeling and…

Computer Vision and Pattern Recognition · Computer Science 2026-03-03 Danfeng Hong , Chenyu Li , Xuyang Li , Gustau Camps-Valls , Jocelyn Chanussot

Existing methods for self-supervised representation learning of geospatial regions and map entities rely extensively on the design of pretext tasks, often involving augmentations or heuristic sampling of positive and negative pairs based on…

Machine Learning · Computer Science 2025-03-11 Theodor Lundqvist , Ludvig Delvret

Medical image grounding aims to align natural language phrases with specific regions in medical images, serving as a foundational task for intelligent diagnosis, visual question answering (VQA), and automated report generation (MRG).…

Computer Vision and Pattern Recognition · Computer Science 2025-11-07 Ziye Deng , Ruihan He , Jiaxiang Liu , Yuan Wang , Zijie Meng , Songtao Jiang , Yong Xie , Zuozhu Liu

While Vision-Language Models (VLMs) have significantly advanced remote sensing interpretation, enabling them to perform complex, step-by-step reasoning remains highly challenging. Recent efforts to introduce Chain-of-Thought (CoT) reasoning…

Computer Vision and Pattern Recognition · Computer Science 2026-05-27 Lang Sun , Ronghao Fu , Zhuoran Duan , Haoran Liu , Xueyan Liu , Bo Yang

Recent open-vocabulary detectors achieve promising performance with abundant region-level annotated data. In this work, we show that an open-vocabulary detector co-training with a large language model by generating image-level detailed…

Computer Vision and Pattern Recognition · Computer Science 2025-02-03 Shenghao Fu , Qize Yang , Qijie Mo , Junkai Yan , Xihan Wei , Jingke Meng , Xiaohua Xie , Wei-Shi Zheng

Multimodal representation learning has been largely driven by contrastive models such as CLIP, which learn a shared embedding space by aligning paired image-text samples. While effective for general-purpose representation learning, such…

Machine Learning · Computer Science 2026-05-12 Yang Qiao , Yuntong Hu , Bowen Zhu , Hasibul Haque , Liang Zhao

Recent progress in large language models (LLMs) has demonstrated the ability to learn and leverage Internet-scale knowledge through pre-training with autoregressive models. Unfortunately, applying such models to settings with embodied…

Visual grounding in 3D is the key for embodied agents to localize language-referred objects in open-world environments. However, existing benchmarks are limited to indoor focus, single-platform constraints, and small scale. We introduce…

Computer Vision and Pattern Recognition · Computer Science 2025-12-02 Rong Li , Yuhao Dong , Tianshuai Hu , Ao Liang , Youquan Liu , Dongyue Lu , Liang Pan , Lingdong Kong , Junwei Liang , Ziwei Liu

Learning generalist embodied agents, able to solve multitudes of tasks in different domains is a long-standing problem. Reinforcement learning (RL) is hard to scale up as it requires a complex reward design for each task. In contrast,…

Artificial Intelligence · Computer Science 2024-11-01 Pietro Mazzaglia , Tim Verbelen , Bart Dhoedt , Aaron Courville , Sai Rajeswar

Remote sensing imagery presents vast, inherently unstructured spatial data, necessitating sophisticated reasoning to interpret complex user intents and contextual relationships beyond simple recognition tasks. In this paper, we aim to…

Computer Vision and Pattern Recognition · Computer Science 2025-12-24 Liang Yao , Fan Liu , Hongbo Lu , Chuanyi Zhang , Rui Min , Shengxiang Xu , Shimin Di , Pai Peng

Building multisensory AI systems that learn from multiple sensory inputs such as text, speech, video, real-world sensors, wearable devices, and medical data holds great promise for impact in many scientific areas with practical benefits,…

Machine Learning · Computer Science 2024-05-01 Paul Pu Liang

Multimodal Entity Linking (MEL) is a crucial task that aims at linking ambiguous mentions within multimodal contexts to the referent entities in a multimodal knowledge base, such as Wikipedia. Existing methods focus heavily on using complex…

Artificial Intelligence · Computer Science 2024-08-22 Liu Qi , He Yongyi , Lian Defu , Zheng Zhi , Xu Tong , Liu Che , Chen Enhong