English

Visual Credit Audit for Multimodal Spatial Reasoning

Computer Vision and Pattern Recognition 2026-07-29 v1 Artificial Intelligence

Abstract

Closed yes/no spatial benchmarks can reward a correct answer even when the image adds little support beyond no-image contexts. Under a fixed forced-choice interface, Visual Credit Audit (VCA) separates two estimands: whether the benchmark image gives the model's declared decision more support than text-only and blank controls, and whether the model responds to relation-specific visual evidence. The first audit is training- and label-free and does not require an answer flip. Applying labels yields dependence-credited correctness (D-CC); on correct items, it equals same-control gold-aligned positive gain, while prediction alignment extends the audit to errors. Across four open MLLMs and two spatial benchmarks, 12.73-26.25% of decisions are correct yet uncredited. Matched same-split image permutation reduces D-CC by 21.25-47.80 points, with every paired 95% interval above zero. Fixed-pixel relation contrasts and a 3x3 evidence-source factorial show why null controls cannot identify relation response. Among controlled correct-but-uncredited agreement decisions, response to relation reversal spans 81.57-100.00%, while 32.11% pooled change answer. Independently audited outcomes on 108 geometry-compatible edits provide a bounded natural-image correspondence check. VCA thereby decomposes benchmark success into correctness, additional image support, and relation-consistent response.

Cite

@article{arxiv.2607.27069,
  title  = {Visual Credit Audit for Multimodal Spatial Reasoning},
  author = {Feixiang Liu and Qiang Qiu and Lanbo Sun and Nan Wei and Huawei Shen and Xueqi Cheng},
  journal= {arXiv preprint arXiv:2607.27069},
  year   = {2026}
}

Comments

"`text 20 pages, 2 figures. Code: https://github.com/SouthWinter/VCA