English

GeoViT: A Versatile Vision Transformer Architecture for Geospatial Image Analysis

Computer Vision and Pattern Recognition 2023-11-27 v1 Machine Learning

Abstract

Greenhouse gases are pivotal drivers of climate change, necessitating precise quantification and source identification to foster mitigation strategies. We introduce GeoViT, a compact vision transformer model adept in processing satellite imagery for multimodal segmentation, classification, and regression tasks targeting CO2 and NO2 emissions. Leveraging GeoViT, we attain superior accuracy in estimating power generation rates, fuel type, plume coverage for CO2, and high-resolution NO2 concentration mapping, surpassing previous state-of-the-art models while significantly reducing model size. GeoViT demonstrates the efficacy of vision transformer architectures in harnessing satellite-derived data for enhanced GHG emission insights, proving instrumental in advancing climate change monitoring and emission regulation efforts globally.

Keywords

Cite

@article{arxiv.2311.14301,
  title  = {GeoViT: A Versatile Vision Transformer Architecture for Geospatial Image Analysis},
  author = {Madhav Khirwar and Ankur Narang},
  journal= {arXiv preprint arXiv:2311.14301},
  year   = {2023}
}

Comments

Extended Abstract, Preprint

R2 v1 2026-06-28T13:30:05.179Z