English

Occlusion Robustness of CLIP for Military Vehicle Classification

Computer Vision and Pattern Recognition 2025-09-03 v2 Artificial Intelligence

Abstract

Vision-language models (VLMs) like CLIP enable zero-shot classification by aligning images and text in a shared embedding space, offering advantages for defense applications with scarce labeled data. However, CLIP's robustness in challenging military environments, with partial occlusion and degraded signal-to-noise ratio (SNR), remains underexplored. We investigate CLIP variants' robustness to occlusion using a custom dataset of 18 military vehicle classes and evaluate using Normalized Area Under the Curve (NAUC) across occlusion percentages. Four key insights emerge: (1) Transformer-based CLIP models consistently outperform CNNs, (2) fine-grained, dispersed occlusions degrade performance more than larger contiguous occlusions, (3) despite improved accuracy, performance of linear-probed models sharply drops at around 35% occlusion, (4) by finetuning the model's backbone, this performance drop occurs at more than 60% occlusion. These results underscore the importance of occlusion-specific augmentations during training and the need for further exploration into patch-level sensitivity and architectural resilience for real-world deployment of CLIP.

Keywords

Cite

@article{arxiv.2508.20760,
  title  = {Occlusion Robustness of CLIP for Military Vehicle Classification},
  author = {Jan Erik van Woerden and Gertjan Burghouts and Lotte Nijskens and Alma M. Liezenga and Sabina van Rooij and Frank Ruis and Hugo J. Kuijf},
  journal= {arXiv preprint arXiv:2508.20760},
  year   = {2025}
}

Comments

To be presented at SPIE: Sensors + Imaging, Artificial Intelligence for Security and Defence Applications II