English

VAD: Attributing Visual Evidence for Target Reconstruction in Multimodal On-Policy Distillation

Computer Vision and Pattern Recognition 2026-07-30 v1 Computation and Language

Abstract

Multimodal on-policy distillation (OPD) transfers fine-grained visual knowledge by supervising student-generated trajectories with a privileged-view teacher. Yet its next-token corrections are source-mixed, combining visual signals with linguistic priors and teacher-specific effects. The key challenge is to estimate which corrections are supported by visual evidence, not merely where or how strongly to distill. We introduce Visual Attribution Distillation (VAD), a counterfactual target-reconstruction algorithm that estimates the visually attributable part of a teacher correction. At each student-generated prefix, VAD evaluates the same fixed teacher with the relevant evidence present and removed. The corresponding change in centered log-probabilities defines ut, a signed proxy for the visual evidence direction that estimates how revealing the evidence supports or refutes candidate tokens. VAD projects the original correction onto this proxy to obtain an intervention-aligned component and a proxy-unexplained residual, then reconstructs a student-anchored target from the former. During training, this reconstructed target supplies the primary supervision signal, while the privileged teacher contributes a weak regularizer. Across six fine-grained visual benchmarks at 4B and 9B scales, VAD outperforms direct privileged-view distillation and visual-advantage weighting. Token- level and controlled-target analyses show that the proxy-aligned component is enriched in task-relevant visual corrections and yields stronger target shifts, especially when evidence refutes a mistaken answer. These results support counterfactual target reconstruction as an effective alternative to source-mixed supervision.

Keywords

Cite

@article{arxiv.2607.28590,
  title  = {VAD: Attributing Visual Evidence for Target Reconstruction in Multimodal On-Policy Distillation},
  author = {Kangning Zhang and Yixing Li and Shuai Shao and Qingyao Li and Zhengxi Lu and Zhiyuan Yao and Jianghao Lin and Wenxiang Jiao and Yuan Lu and Weiwen Liu and Weinan Zhang and Yong Yu},
  journal= {arXiv preprint arXiv:2607.28590},
  year   = {2026}
}

Comments

The project is accessible at https://github.com/DeepExperience/VAD_Multimodal_OPD