中文

基于理由引导的知识图书情报视觉问答

计算与语言 2025-08-08 v3 人工智能

摘要

最近,大型语言模型(LLM)被用于知识图书情报视觉问答(VQA)。尽管先前研究取得了令人鼓舞的结果,但先前的方法只提示LLM直接预测答案,忽略了中间思考过程。我们认为先前的方法未充分激发LLM的潜能。我们提出了一个名为PLRH的框架,该框架为知识图书情报VQA中的LLM提供理由启发式提示(PLRH that Prompts LLMs with Rationale Heuristics for knowledge-based VQA)。PLRH提示LLM生成理由启发式,即中间思考过程,然后利用理由启发式激发LLM预测答案。实验表明,我们的方法在OK-VQA和A-OKVQA上的表现分别超过了现有基线2.2和2.1。

关键词

引用

@article{arxiv.2412.16936,
  title  = {Rationale-guided Prompting for Knowledge-based Visual Question Answering},
  author = {Zhongjian Hu and Peng Yang and Bing Li and Fengyuan Liu},
  journal= {arXiv preprint arXiv:2412.16936},
  year   = {2025}
}

备注

We would like to withdraw this submission due to ongoing internal review and coordination among the author team. Upon the supervisor's recommendation, we have decided to delay public dissemination until the manuscript undergoes further refinement and aligns with our intended academic trajectory