See the Forest and the Trees: 面向知识库视觉问答的协同推理框架
计算机视觉与模式识别
2025-08-14 v3
摘要
多模态大语言模型(MLLM)已推动了知识库视觉问答(KBVQA)领域的前沿进展,但其推理本质上受限于对单一维度证据的依赖。这种“只见树木,不见森林”的方法阻碍了健壮、多维度的理解。以“既见森林又见树木”为原则,我们提出了Synergos-VQA,一种新的协同推理框架。我们的框架在推理时同时生成并融合三种互补证据流:(1)整体证据以感知整个场景(“森林”),(2)来自原型驱动模块的结构证据以识别关键对象(“树木”),以及(3)来自反事实探针的因果证据以确保推理的稳健性。通过协同融合多维证据,我们的框架实现了更全面可靠的推理过程。大量实验表明,Synergos-VQA在三个具有挑战性的基准测试上(包括OK-VQA和A-OKVQA)上明确建立了新的状态记录。此外,我们的方法显示出强大的即插即用能力,显著提升了各种开源MLLM的性能,证明了卓越的方法学设计可以超过单纯的模型规模。
引用
@article{arxiv.2507.17659,
title = {See the Forest and the Trees: A Synergistic Reasoning Framework for Knowledge-Based Visual Question Answering},
author = {Junjie Wang and Yunhan Tang and Yijie Wang and Zhihao Yuan and Huan Wang and Yangfan He and Bin Li},
journal= {arXiv preprint arXiv:2507.17659},
year = {2025}
}
备注
We are withdrawing this preprint because it is undergoing a major revision and restructuring. We feel that the current version does not convey our core contributions and methodology with sufficient clarity and accuracy