English

ACD-CLIP: Decoupling Representation and Dynamic Fusion for Zero-Shot Anomaly Detection

Computer Vision and Pattern Recognition 2026-03-30 v6 Artificial Intelligence Machine Learning

Abstract

Pre-trained Vision-Language Models (VLMs) struggle with Zero-Shot Anomaly Detection (ZSAD) due to a critical adaptation gap: they lack the local inductive biases required for dense prediction and employ inflexible feature fusion paradigms. We address these limitations through an Architectural Co-Design framework that jointly refines feature representation and cross-modal fusion. Our method proposes a parameter-efficient Convolutional Low-Rank Adaptation (Conv-LoRA) adapter to inject local inductive biases for fine-grained representation, and introduces a Dynamic Fusion Gateway (DFG) that leverages visual context to adaptively modulate text prompts, enabling a powerful bidirectional fusion. Extensive experiments on diverse industrial and medical benchmarks demonstrate superior accuracy and robustness, validating that this synergistic co-design is critical for robustly adapting foundation models to dense perception tasks. The source code is available at https://github.com/cockmake/ACD-CLIP.

Keywords

Cite

@article{arxiv.2508.07819,
  title  = {ACD-CLIP: Decoupling Representation and Dynamic Fusion for Zero-Shot Anomaly Detection},
  author = {Ke Ma and Jun Long and Hongxiao Fei and Liujie Hua and Zhen Dai and Yueyi Luo},
  journal= {arXiv preprint arXiv:2508.07819},
  year   = {2026}
}

Comments

4 pages, 1 reference, 3 figures

R2 v1 2026-07-01T04:43:59.404Z