English

VT-ADL: A Vision Transformer Network for Image Anomaly Detection and Localization

Computer Vision and Pattern Recognition 2021-11-03 v1 Artificial Intelligence Machine Learning

Abstract

We present a transformer-based image anomaly detection and localization network. Our proposed model is a combination of a reconstruction-based approach and patch embedding. The use of transformer networks helps to preserve the spatial information of the embedded patches, which are later processed by a Gaussian mixture density network to localize the anomalous areas. In addition, we also publish BTAD, a real-world industrial anomaly dataset. Our results are compared with other state-of-the-art algorithms using publicly available datasets like MNIST and MVTec.

Keywords

Cite

@article{arxiv.2104.10036,
  title  = {VT-ADL: A Vision Transformer Network for Image Anomaly Detection and Localization},
  author = {Pankaj Mishra and Riccardo Verk and Daniele Fornasier and Claudio Piciarelli and Gian Luca Foresti},
  journal= {arXiv preprint arXiv:2104.10036},
  year   = {2021}
}

Comments

6 Pages, 4 images, conference published paper

R2 v1 2026-06-24T01:22:20.655Z