中文

面向文档视觉问答的视觉语言模型空间导向解释

计算机视觉与模式识别 2025-07-18 v1 人工智能 计算与语言 机器学习

摘要

我们提出 EaGERS,一个完全训练-free且模型无关的管道, (1) 通过视觉语言模型生成自然语言理由, (2) 通过在可配置网格上计算多模态嵌入相似性并进行多数投票,将这些理由对齐到空间子区域, (3) 仅从相关区域(即被遮盖的图像)中限制响应的生成。在 DocVQA 数据集上的实验表明,我们最佳配置不仅在 exact match accuracy 和 Average Normalized Levenshtein Similarity 指标上超越基础模型,还在无需额外模型微调的情况下增强了 DocVQA 的透明度和可复现性。

关键词

引用

@article{arxiv.2507.12490,
  title  = {Spatially Grounded Explanations in Vision Language Models for Document Visual Question Answering},
  author = {Maximiliano Hormazábal Lagos and Héctor Cerezo-Costas and Dimosthenis Karatzas},
  journal= {arXiv preprint arXiv:2507.12490},
  year   = {2025}
}

备注

This work has been accepted for presentation at the 16th Conference and Labs of the Evaluation Forum (CLEF 2025) and will be published in the proceedings by Springer in the Lecture Notes in Computer Science (LNCS) series. Please cite the published version when available