English

VID-AD: A Dataset for Image-Level Logical Anomaly Detection under Vision-Induced Distraction

Computer Vision and Pattern Recognition 2026-03-17 v1

Abstract

Logical anomaly detection in industrial inspection remains challenging due to variations in visual appearance (e.g., background clutter, illumination shift, and blur), which often distract vision-centric detectors from identifying rule-level violations. However, existing benchmarks rarely provide controlled settings where logical states are fixed while such nuisance factors vary. To address this gap, we introduce VID-AD, a dataset for logical anomaly detection under vision-induced distraction. It comprises 10 manufacturing scenarios and five capture conditions, totaling 50 one-class tasks and 10,395 images. Each scenario is defined by two logical constraints selected from quantity, length, type, placement, and relation, with anomalies including both single-constraint and combined violations. We further propose a language-based anomaly detection framework that relies solely on text descriptions generated from normal images. Using contrastive learning with positive texts and contradiction-based negative texts synthesized from these descriptions, our method learns embeddings that capture logical attributes rather than low-level features. Extensive experiments demonstrate consistent improvements over baselines across the evaluated settings. The dataset is available at: https://github.com/nkthiroto/VID-AD.

Keywords

Cite

@article{arxiv.2603.13964,
  title  = {VID-AD: A Dataset for Image-Level Logical Anomaly Detection under Vision-Induced Distraction},
  author = {Hiroto Nakata and Yawen Zou and Shunsuke Sakai and Shun Maeda and Chunzhi Gu and Yijin Wei and Shangce Gao and Chao Zhang},
  journal= {arXiv preprint arXiv:2603.13964},
  year   = {2026}
}
R2 v1 2026-07-01T11:20:06.356Z