English

MMLGNet: Cross-Modal Alignment of Remote Sensing Data using CLIP

Computer Vision and Pattern Recognition 2026-01-14 v1

Abstract

In this paper, we propose a novel multimodal framework, Multimodal Language-Guided Network (MMLGNet), to align heterogeneous remote sensing modalities like Hyperspectral Imaging (HSI) and LiDAR with natural language semantics using vision-language models such as CLIP. With the increasing availability of multimodal Earth observation data, there is a growing need for methods that effectively fuse spectral, spatial, and geometric information while enabling semantic-level understanding. MMLGNet employs modality-specific encoders and aligns visual features with handcrafted textual embeddings in a shared latent space via bi-directional contrastive learning. Inspired by CLIP's training paradigm, our approach bridges the gap between high-dimensional remote sensing data and language-guided interpretation. Notably, MMLGNet achieves strong performance with simple CNN-based encoders, outperforming several established multimodal visual-only methods on two benchmark datasets, demonstrating the significant benefit of language supervision. Codes are available at https://github.com/AdityaChaudhary2913/CLIP_HSI.

Keywords

Cite

@article{arxiv.2601.08420,
  title  = {MMLGNet: Cross-Modal Alignment of Remote Sensing Data using CLIP},
  author = {Aditya Chaudhary and Sneha Barman and Mainak Singha and Ankit Jha and Girish Mishra and Biplab Banerjee},
  journal= {arXiv preprint arXiv:2601.08420},
  year   = {2026}
}

Comments

Accepted at InGARSS 2025

R2 v1 2026-07-01T09:02:32.567Z