用于卫星图像时间序列作物映射的多模态视觉Transformer
计算机视觉与模式识别
2024-06-25 v1
摘要
使用由不同卫星传感器获取的图像在卫星图像时间序列(SITS)作物映射框架中已显示可提高分类性能。现有最先进的体系结构使用自注意力机制处理SITS的时间维度,并使用卷积处理空间维度。鼓励纯注意力体系结构在单模态SITS作物映射中取得的成功,我们引入了几种基于Transformer的多模态多时序体系结构。具体而言,我们研究了在Temporo-Spatial Vision Transformer(TSViT)中Early Fusion、Cross Attention Fusion和Synchronized Class Token Fusion的有效性。实验结果表明,实验结果在包含卷积和自注意力组件的体系结构方面均取得了显著的改进。
引用
@article{arxiv.2406.16513,
title = {Multi-Modal Vision Transformers for Crop Mapping from Satellite Image Time Series},
author = {Theresa Follath and David Mickisch and Jan Hemmerling and Stefan Erasmi and Marcel Schwieder and Begüm Demir},
journal= {arXiv preprint arXiv:2406.16513},
year = {2024}
}
备注
5 pages, 2 figures, 1 table. Accepted at IEEE International Geoscience and Remote Sensing Symposium (IGARSS) 2024. Our code is available at https://git.tu-berlin.de/rsim/mmtsvit