English

Synergistic Perception and Generative Recomposition: A Multi-Agent Orchestration for Expert-Level Building Inspection

Computer Vision and Pattern Recognition 2026-05-05 v2

Abstract

Building facade defect inspection is fundamental to structural health monitoring and sustainable urban maintenance, yet it remains a formidable challenge due to extreme geometric variability, low contrast against complex backgrounds, and the inherent complexity of composite defects (e.g., cracks co-occurring with spalling). Such characteristics lead to severe pixel imbalance and feature ambiguity, which, coupled with the critical scarcity of high-quality pixel-level annotations, hinder the generalization of existing detection and segmentation models. To address gaps, we propose \textit{FacadeFixer}, a unified multi-agent framework that treats defect perception as a collaborative reasoning task rather than isolated recognition. Specifically,\textit{FacadeFixer} orchestrates specialized agents for detection and segmentation to handle multi-type defect interference, working in tandem with a generative agent to enable semantic recomposition. This process decouples intricate defects from noisy backgrounds and realistically synthesizes them onto diverse clean textures, generating high-fidelity augmented data with precise expert-level masks. To support this, we introduce a comprehensive multi-task dataset covering six primary facade categories with pixel-level annotations. Extensive experiments demonstrate that \textit{FacadeFixer} significantly outperforms state-of-the-art (SOTA) baselines. Specifically, it excels in capturing pixel-level structural anomalies and highlights generative synthesis as a robust solution to data scarcity in infrastructure inspection. Our code and dataset will be made publicly available.

Keywords

Cite

@article{arxiv.2603.20143,
  title  = {Synergistic Perception and Generative Recomposition: A Multi-Agent Orchestration for Expert-Level Building Inspection},
  author = {Hui Zhong and Yichun Gao and Luyan Liu and Xusen Guo and Zhaonian Kuang and Qiming Zhang and Xinhu Zheng},
  journal= {arXiv preprint arXiv:2603.20143},
  year   = {2026}
}

Comments

We are withdrawing this article because we recently identified a major methodological error regarding the multi-agent orchestration setup described in Section 4.2. This issue significantly impacts the final conclusions drawn in the paper. We sincerely apologize for any confusion this may have caused

R2 v1 2026-07-01T11:30:05.749Z