English

C-DiffDet+: Fusing Global Scene Context with Generative Denoising for High-Fidelity Car Damage Detection

Computer Vision and Pattern Recognition 2025-10-28 v4

Abstract

Fine-grained object detection in challenging visual domains, such as vehicle damage assessment, presents a formidable challenge even for human experts to resolve reliably. While DiffusionDet has advanced the state-of-the-art through conditional denoising diffusion, its performance remains limited by local feature conditioning in context-dependent scenarios. We address this fundamental limitation by introducing Context-Aware Fusion (CAF), which leverages cross-attention mechanisms to integrate global scene context with local proposal features directly. The global context is generated using a separate dedicated encoder that captures comprehensive environmental information, enabling each object proposal to attend to scene-level understanding. Our framework significantly enhances the generative detection paradigm by enabling each object proposal to attend to comprehensive environmental information. Experimental results demonstrate an improvement over state-of-the-art models on the CarDD benchmark, establishing new performance benchmarks for context-aware object detection in fine-grained domains

Keywords

Cite

@article{arxiv.2509.00578,
  title  = {C-DiffDet+: Fusing Global Scene Context with Generative Denoising for High-Fidelity Car Damage Detection},
  author = {Abdellah Zakaria Sellam and Ilyes Benaissa and Salah Eddine Bekhouche and Abdenour Hadid and Vito Renó and Cosimo Distante},
  journal= {arXiv preprint arXiv:2509.00578},
  year   = {2025}
}
R2 v1 2026-07-01T05:13:38.541Z