中文

面向鲁棒图像字幕评估的分布感知评分解码器 DISCODE

计算机视觉与模式识别 2026-01-06 v2 人工智能

摘要

大型视觉语言模型(LVLMs)在广泛的多模态任务中展现出惊人的性能。然而,在 domain-shift 场景下使用 LVLMs 对图像字幕进行鲁棒评估仍具有挑战性。为此,我们提出 Distribution-Aware Score Decoder(DISCODE),一种 novel finetuning-free 方法,用于生成在 diverse 域之间更符合人类判断的鲁棒评分。DISCODE 的核心思想是其 test-time adaptive evaluation 方法,引入 Adaptive Test-Time(ATT)loss,利用高斯先验分布提高评分估计的鲁棒性。该 loss 通过我们推导的解析解在 test time 有效地最小化。进一步,我们引入 Multi-domain Caption Evaluation(MCEval)基准,一个覆盖六个 distinct 域的新的图像字幕评估基准,用于评估评估指标的鲁棒性。在实验中,我们展示 DISCODE 在 MCEval 和四个代表性现有基准上作为 reference-free 评估指标实现了最佳性能。

关键词

引用

@article{arxiv.2512.14420,
  title  = {DISCODE: Distribution-Aware Score Decoder for Robust Automatic Evaluation of Image Captioning},
  author = {Nakamasa Inoue and Kanoko Goto and Masanari Oi and Martyna Gruszka and Mahiro Ukai and Takumi Hirose and Yusuke Sekikawa},
  journal= {arXiv preprint arXiv:2512.14420},
  year   = {2026}
}

备注

Paper accepted to AAAI 2026