中文
相关论文

相关论文: FUSAR-KLIP: Towards Multimodal Foundation Models f…

200 篇论文

Image captioning has become an important task in computer vision, enabling models to generate natural language descriptions of visual content. While several datasets exist for natural images and high-resolution optical remote sensing…

计算机视觉与模式识别 · 计算机科学 2026-05-06 Lucrezia Tosato , Gianluca Lombardi , Ronny Hansch

Fine-grained vision-language understanding requires precise alignment between visual content and linguistic descriptions, a capability that remains limited in current models, particularly in non-English settings. While models like CLIP…

计算机视觉与模式识别 · 计算机科学 2026-05-26 Chunyu Xie , Bin Wang , Fanjing Kong , Jincheng Li , Dawei Liang , Ji Ao , Dawei Leng , Yuhui Yin

Remote Sensing Image Captioning (RSIC) is a cross-modal field bridging vision and language, aimed at automatically generating natural language descriptions of features and scenes in remote sensing imagery. Despite significant advances in…

计算机视觉与模式识别 · 计算机科学 2025-03-07 Qing Zhou , Tao Yang , Junyu Gao , Weiping Ni , Junzheng Wu , Qi Wang

As remote sensing (RS) data obtained from different sensors become available largely and openly, multimodal data processing and analysis techniques have been garnering increasing interest in the RS and geoscience community. However, due to…

计算机视觉与模式识别 · 计算机科学 2021-05-24 Danfeng Hong , Jingliang Hu , Jing Yao , Jocelyn Chanussot , Xiao Xiang Zhu

In geographical image segmentation, performance is often constrained by the limited availability of training data and a lack of generalizability, particularly for segmenting mobility infrastructure such as roads, sidewalks, and crosswalks.…

计算机视觉与模式识别 · 计算机科学 2025-08-19 Rafi Ibn Sultan , Chengyin Li , Hui Zhu , Prashant Khanduri , Marco Brocanelli , Dongxiao Zhu

Vision-language models like CLIP have shown impressive capabilities in aligning images and text, but they often struggle with lengthy and detailed text descriptions because of their training focus on short and concise captions. We present…

计算机视觉与模式识别 · 计算机科学 2025-03-26 Hyungyu Choi , Young Kyun Jang , Chanho Eom

Prior studies on Remote Sensing Foundation Model (RSFM) reveal immense potential towards a generic model for Earth Observation. Nevertheless, these works primarily focus on a single modality without temporal and geo-context modeling,…

In this work we investigate the viability of foundational AI/ML models for Synthetic Aperture Radar (SAR) object recognition tasks. We are inspired by the tremendous progress being made in the wider community, particularly in the natural…

计算机视觉与模式识别 · 计算机科学 2025-09-29 Nathan Inkawhich

Recent advances in multimodal large language models (MLLMs) have enabled impressive progress in vision-language understanding, yet their high computational cost limits deployment in resource-constrained scenarios such as robotic…

计算机视觉与模式识别 · 计算机科学 2025-11-26 Quoc-Huy Trinh

The growing availability of co-located geospatial data spanning aerial imagery, street-level views, elevation models, text, and geographic coordinates offers a unique opportunity for multimodal representation learning. We introduce…

计算机视觉与模式识别 · 计算机科学 2026-04-14 Guillaume Astruc , Eduard Trulls , Jan Hosang , Loic Landrieu , Paul-Edouard Sarlin

Synthetic Aperture Radar (SAR) imaging results are highly sensitive to observation geometries and the geometric parameters of targets. However, existing generative methods primarily operate within the image domain, neglecting explicit…

图像与视频处理 · 电气工程与系统科学 2026-01-08 Fan Zhang , Xuanting Wu , Fei Ma , Qiang Yin , Yuxin Hu

Recent advancements in foundation models have improved autonomous tool usage and reasoning, but their capabilities in map-based reasoning remain underexplored. To address this, we introduce MapEval, a benchmark designed to assess foundation…

Optical-SAR image matching is a fundamental task for image fusion and visual navigation. However, all large-scale open SAR dataset for methods development are collected from single platform, resulting in limited satellite types and spatial…

计算机视觉与模式识别 · 计算机科学 2025-10-14 Yibin Ye , Xichao Teng , Shuo Chen , Yijie Bian , Tao Tan , Zhang Li

Fine-grained image-text alignment is a pivotal challenge in multimodal learning, underpinning key applications such as visual question answering, image captioning, and vision-language navigation. Unlike global alignment, fine-grained…

计算机视觉与模式识别 · 计算机科学 2025-12-02 Jiale Liu , Haoming Zhou , Yishu Liu , Bingzhi Chen , Yuncheng Jiang

In this paper, we propose a novel multimodal framework, Multimodal Language-Guided Network (MMLGNet), to align heterogeneous remote sensing modalities like Hyperspectral Imaging (HSI) and LiDAR with natural language semantics using…

计算机视觉与模式识别 · 计算机科学 2026-01-14 Aditya Chaudhary , Sneha Barman , Mainak Singha , Ankit Jha , Girish Mishra , Biplab Banerjee

The application of Vision-Language Models (VLMs) in remote sensing (RS) image understanding has achieved notable progress, demonstrating the basic ability to recognize and describe geographical entities. However, existing RS-VLMs are mostly…

计算机视觉与模式识别 · 计算机科学 2025-07-21 Xianzhi Ma , Jianhui Li , Changhua Pei , Hao Liu

Geo-spatial analysis of our world benefits from a multimodal approach, as every single geographic location can be described in numerous ways (images from various viewpoints, textual descriptions, geographic coordinates, etc.). Current…

计算机视觉与模式识别 · 计算机科学 2026-04-29 Oskar Kristoffersen , Alba Reinders Sánchez , Morten Rieger Hannemose , Anders Bjorholm Dahl , Dim P. Papadopoulos

Multimodal Large Language Models (MLLMs) have achieved remarkable success in vision-language tasks but their remote sensing (RS) counterpart are relatively under explored. Unlike natural images, RS imagery presents unique challenges that…

计算机视觉与模式识别 · 计算机科学 2025-05-13 Abduljaleel Adejumo , Faegheh Yeganli , Clifford Broni-bediako , Aoran Xiao , Naoto Yokoya , Mennatullah Siam

Deep Learning (DL) is undergoing a paradigm shift with the emergence of foundation models. In this work, we focus on Contrastive Language-Image Pre-training (CLIP), a Vision-Language foundation model that achieves high accuracy across…

计算机视觉与模式识别 · 计算机科学 2025-07-21 Angelos Zavras , Dimitrios Michail , Begüm Demir , Ioannis Papoutsis

Multi-modal large language models (MLLMs) have demonstrated remarkable success in vision and visual-language tasks within the natural image domain. Owing to the significant diversities between the natural and remote sensing (RS) images, the…

计算机视觉与模式识别 · 计算机科学 2024-03-11 Wei Zhang , Miaoxin Cai , Tong Zhang , Yin Zhuang , Xuerui Mao