中文

MM-OPERA: 用于大型视觉语言模型的开放式关联推理基准

组合数学 2025-11-03 v1

摘要

大型视觉语言模型(LVLMs)已展现出显著进步。然而,与 human intelligence 相比仍存在不足,如 hallucination 和浅层模式匹配。在本 work 中,我们旨在评估一种 fundamental 且尚未充分探索的 intelligence:association,是 human cognition 用于 creative thinking 和 knowledge integration 的基石。当前的 benchmark 常受限于 closed-ended tasks,无法捕捉至关重要于 real-world applications 的 open-ended association reasoning 的复杂性。为此,我们提出 MM-OPERA,一个 systematical benchmark,包含 11,497 个实例,涵盖 two open-ended tasks: Remote-Item Association (RIA) 和 In-Context Association (ICA),将 association intelligence 评估与 human psychometric 原理相结合。它挑战 LVLMs 通过 free-form responses 和 explicit reasoning paths 来模拟 divergent thinking 和 convergent associative reasoning 的 spirit。我们部署 tailored LLM-as-a-Judge 策略以评估 open-ended outputs,应用 process-reward-informed judgment 对 reasoning 进行精细分析。广泛的 empirical 研究涉及 state-of-the-art LVLMs,包括 task instance 的 sensitivity analysis、LLM-as-a-Judge 策略的 validity analysis 以及 abilities、domains、languages、cultures 等方面的 diversity analysis,为全面细致地理解 current LVLMs 在 associative reasoning 方面的局限性提供了依据,为更贴近 human-like 和 general-purpose AI 的道路铺平了轼道。该 dataset 和 code 可在 https://github.com/MM-OPERA-Bench/MM-OPERA 上获取。

关键词

引用

@article{arxiv.2510.26936,
  title  = {Comparing the numbers of subforests and subgraph-degree-tuples},
  author = {Sergei Shteiner and Pavel Shteyner},
  journal= {arXiv preprint arXiv:2510.26936},
  year   = {2025}
}

备注

24 pages