English

Seen, Said, or Forgotten? A Causal Audit of Visual KV Memory Across Dialog Turns

Computer Vision and Pattern Recognition 2026-07-28 v1 Artificial Intelligence

Abstract

Stateful multimodal assistants encode an image once but may answer questions about it many turns later. Attention-guided visual-KV eviction assumes that evidence irrelevant now will remain dispensable, although future questions are unknown. We ask when a visual fact is actually safe to forget and introduce the Causal Visual Memory Audit (CVMA), a paired single-prefill framework that tests what later answers lose when a visual region, the whole image, or prior assistant text becomes unavailable. On VisDial and ConvBench, current attention can rank future-useful regions worse than random even though a diagnostic marginal-utility control shows substantial selection headroom. Aggregate scores hide this failure when later turns do not need vision; controlled and stock-generated histories reveal a second escape route, in which assistant-text KV replaces image KV for facts already stated but not reliably for unstated facts. In the tested stacks, safe forgetting is supported by low future visual dependence or fact-specific verbalization---not by low current attention.

Keywords

Cite

@article{arxiv.2607.25467,
  title  = {Seen, Said, or Forgotten? A Causal Audit of Visual KV Memory Across Dialog Turns},
  author = {Hong Chen and Kang Chen and Yuxuan Fan and Bo Wang and Yubo Gao and Yuanlin Chu and Xuming Hu},
  journal= {arXiv preprint arXiv:2607.25467},
  year   = {2026}
}