English

VEGAS: Human-Aligned Video Caption Evaluation via Gaze

Computer Vision and Pattern Recognition 2026-07-09 v1 Artificial Intelligence Human-Computer Interaction

Abstract

Vision-language models excel at video captioning, yet typically generate descriptions that fail to capture individual viewers' attention. We propose VEGAS (Video caption Evaluation via GAze Score), a training-free metric that leverages test-time gaze to sample personalized, attention-aligned text. It is a cross-modal, information-theoretic metric that quantifies how well a candidate caption matches a viewer's focus. To evaluate VEGAS, we curate a dataset of egocentric activities and instructional slides paired with synchronized gaze and reference annotations. We then select captions based on VEGAS via rejection sampling without model retraining. Experiments show that VEGAS-selected captions align significantly better with human focus and improve downstream caption-to-video retrieval, demonstrating the practical utility of incorporating viewer attention during inference.

Cite

@article{arxiv.2607.08489,
  title  = {VEGAS: Human-Aligned Video Caption Evaluation via Gaze},
  author = {Shenghui Chen and Po-han Li and Ximeng Sun and Shijia Yang and Emad Barsoum and Zicheng Liu and Sandeep Chinchali and Ufuk Topcu},
  journal= {arXiv preprint arXiv:2607.08489},
  year   = {2026}
}
R2 v1 2026-07-22T20:32:27.651Z