English

DFM: Difference Feature Modeling with Text-Guided Gated Contrastive Loss for Remote Sensing Image Change Captioning

Image and Video Processing 2026-06-25 v1 Machine Learning

Abstract

The primary goal of Remote Sensing Image Change Captioning (RSICC) is to automatically generate descriptions of changes between remote sensing images captured at different time points. Existing models still rely on a single autoregressive generation paradigm, which tends to prioritize learning easily generated vocabulary over capturing discriminative differences between images. To address this, we reframe the training paradigm and propose a novel Difference Feature Modeling (DFM) framework. Specifically, we introduce a Text-guided Gated Contrastive Loss (TGCL) to guide the vision encoder to extract critical features from a text-modal perspective. Additionally, we incorporate a pre-trained Change Detection model to transfer stable change detection knowledge. In order to further enhance the representation, we design a Joint Feature Modeling (JFM) module to achieve the fusion of multi-scale difference representations, thereby capturing comprehensive spatiotemporal variations between multi-temporal images. Extensive experiments on multiple datasets demonstrate the effectiveness of our approach.

Cite

@article{arxiv.2606.27410,
  title  = {DFM: Difference Feature Modeling with Text-Guided Gated Contrastive Loss for Remote Sensing Image Change Captioning},
  author = {Yelin Wang and Zijia Song and Chuanguang Yang and Miaoyu Wang and Zhulin An and Libo Huang and Yongjun Xu},
  journal= {arXiv preprint arXiv:2606.27410},
  year   = {2026}
}

Comments

Accepted by IEEE ICME 2026

R2 v1 2026-07-22T20:10:47.689Z