English

SEMT: Static-Expansion-Mesh Transformer Network Architecture for Remote Sensing Image Captioning

Computer Vision and Pattern Recognition 2025-07-18 v1 Artificial Intelligence

Abstract

Image captioning has emerged as a crucial task in the intersection of computer vision and natural language processing, enabling automated generation of descriptive text from visual content. In the context of remote sensing, image captioning plays a significant role in interpreting vast and complex satellite imagery, aiding applications such as environmental monitoring, disaster assessment, and urban planning. This motivates us, in this paper, to present a transformer based network architecture for remote sensing image captioning (RSIC) in which multiple techniques of Static Expansion, Memory-Augmented Self-Attention, Mesh Transformer are evaluated and integrated. We evaluate our proposed models using two benchmark remote sensing image datasets of UCM-Caption and NWPU-Caption. Our best model outperforms the state-of-the-art systems on most of evaluation metrics, which demonstrates potential to apply for real-life remote sensing image systems.

Keywords

Cite

@article{arxiv.2507.12845,
  title  = {SEMT: Static-Expansion-Mesh Transformer Network Architecture for Remote Sensing Image Captioning},
  author = {Khang Truong and Lam Pham and Hieu Tang and Jasmin Lampert and Martin Boyer and Son Phan and Truong Nguyen},
  journal= {arXiv preprint arXiv:2507.12845},
  year   = {2025}
}
R2 v1 2026-07-01T04:05:34.076Z