English

How Many Visual Levers Drive Urban Perception? Interventional Counterfactuals via Multiple Localised Edits

Computers and Society 2026-04-27 v1 Computer Vision and Pattern Recognition

Abstract

Street-view perception models predict subjective attributes such as safety at scale, but remain correlational: they do not identify which localized visual changes would plausibly shift human judgement for a specific scene. We propose a lever-based interventional counterfactual framework that recasts scene-level explainability as a bounded search over structured counterfactual edits. Each lever specifies a semantic concept, spatial support, intervention direction, and constrained edit template. Candidate edits are generated through prompt-conditioned image editing and retained only if they satisfy validity checks for same-place preservation, locality, realism, and plausibility. In a pilot across 50 scenes from five cities, the framework reveals preliminary proxy-based directional patterns and a practical failure taxonomy under prompt-only editing, with Mobility Infrastructure and Physical Maintenance showing the largest auxiliary safety shifts. Human pairwise judgements remain the ground-truth endpoint for future validation.

Keywords

Cite

@article{arxiv.2604.22103,
  title  = {How Many Visual Levers Drive Urban Perception? Interventional Counterfactuals via Multiple Localised Edits},
  author = {Jason Tang and Stephen Law},
  journal= {arXiv preprint arXiv:2604.22103},
  year   = {2026}
}
R2 v1 2026-07-01T12:33:09.408Z