English

Explanation-Driven Counterfactual Testing for Faithfulness in Vision-Language Model Explanations

Computer Vision and Pattern Recognition 2025-10-02 v1 Artificial Intelligence

Abstract

Vision-Language Models (VLMs) often produce fluent Natural Language Explanations (NLEs) that sound convincing but may not reflect the causal factors driving predictions. This mismatch of plausibility and faithfulness poses technical and governance risks. We introduce Explanation-Driven Counterfactual Testing (EDCT), a fully automated verification procedure for a target VLM that treats the model's own explanation as a falsifiable hypothesis. Given an image-question pair, EDCT: (1) obtains the model's answer and NLE, (2) parses the NLE into testable visual concepts, (3) generates targeted counterfactual edits via generative inpainting, and (4) computes a Counterfactual Consistency Score (CCS) using LLM-assisted analysis of changes in both answers and explanations. Across 120 curated OK-VQA examples and multiple VLMs, EDCT uncovers substantial faithfulness gaps and provides regulator-aligned audit artifacts indicating when cited concepts fail causal tests.

Keywords

Cite

@article{arxiv.2510.00047,
  title  = {Explanation-Driven Counterfactual Testing for Faithfulness in Vision-Language Model Explanations},
  author = {Sihao Ding and Santosh Vasa and Aditi Ramadwar},
  journal= {arXiv preprint arXiv:2510.00047},
  year   = {2025}
}

Comments

NeurIPS 2025 workshop on Regulatable ML

R2 v1 2026-07-01T06:08:35.424Z