English

Spatial-Frequency Gated Swin Transformer for Remote Sensing Single-Image Super-Resolution

Computer Vision and Pattern Recognition 2026-05-12 v1

Abstract

Remote Sensing (RS) single-image super-resolution aims to reconstruct high-resolution imagery from low-resolution observations while preserving fine spatial structures. Recent Swin Transformer-based models, including Swin2SR, provide strong spatial context modeling throughshifted-window self-attention, but their feed-forward networks remain generic channel-mixing modules and do not separate low-frequency structural content from high-frequency residual detail. To address this limitation, we propose SFG-SwinSR, a Spatial-Frequency Gated Swin Transformer for single-image super-resolution in remote sensing. SFG-SwinSR modifies the original Swin2SR attention block by replacing each transformer block's standard feed-forward network with a lightweight Spatial-Frequency Gated Feed-Forward Network (SFG-FFN). The module estimates low-frequency content via a depthwise-blur branch, extracts high-frequency residuals by subtraction, refines them with a lightweight spatial branch, and adaptively injects detail through a bottleneck gate. Experiments on SpaceNet and SEN2VEN{\mu}S show that SFG-SwinSR improves reconstruction quality under the evaluated settings. On SpaceNet, it achieves 45.19 dB PSNR and 0.9852 SSIM, indicating effective enhancement of high-frequency details. This demonstrates that spatial-frequency transformation within the transformer feed-forward network improves detail reconstruction in RS super-resolution.

Keywords

Cite

@article{arxiv.2605.09687,
  title  = {Spatial-Frequency Gated Swin Transformer for Remote Sensing Single-Image Super-Resolution},
  author = {Md Aminur Hossain and Parekh Valkesh and Ayush V. Patel and Yogesh Jethani and Sanjay K. Singh and Biplab Banerjee},
  journal= {arXiv preprint arXiv:2605.09687},
  year   = {2026}
}

Comments

15 pages

R2 v1 2026-07-22T07:02:33.633Z