English

ViSAGE: Constructing Self-Correcting Memories for Long-Form Video Understanding

Artificial Intelligence 2026-07-29 v1 Computer Vision and Pattern Recognition

Abstract

Multimodal agents operating in long-horizon environments must build and continually update multimedia memories to support entity-consistent, temporally grounded reasoning. However, existing agentic memory approaches often discard fine-grained dentity cues under aggressive compression and segment-wise processing. They also rely heavily on vector similarity retrieval, which can surface semantically related yet identity-mismatched evidence, leading to entity confusion, error propagation, and hallucinated answers. We propose ViSAGE, a multimodal agentic memory framework that constructs self-correcting, entity-centric memories. Specifically, ViSAGE anchors entity identity via cross-modal binding over long temporal ranges. It then applies bidirectional memory refinement to propagate delayed identity evidence, retroactively unifying historical records and improving future reasoning. We also introduce multi-agent cross-verification to assess retrieved evidence under an identity-evidence alignment onstraint, enabling abstention instead of unsupported answers when evidence is missing. Extensive results demonstrate that ViSAGE consistently outperforms the strongest baseline, achieving 5.9% higher accuracy.

Cite

@article{arxiv.2607.28678,
  title  = {ViSAGE: Constructing Self-Correcting Memories for Long-Form Video Understanding},
  author = {Xinkui Zhao and Enbo Chen and Yifan Zhang and Chang Liu and Guanjie Cheng and Naibo Wang and Yueshen Xu},
  journal= {arXiv preprint arXiv:2607.28678},
  year   = {2026}
}

Comments

Accept by ACMMM 2026