MoNaCo:面向数十文档推理的更自然且更复杂的问题基准
计算与语言
2025-09-04 v2 人工智能
数据库
摘要
由大型语言模型(LLM)驱动的自动化代理正在成为查询信息的首选工具。然而,LLM代理的评估基准很少包含既信息寻求又对人类真正耗时的自然问题。为填补这一差距,我们引入MoNaCo,这是一个包含1,315个自然且耗时的questions的数据集,这些questions需要解决数十个乃至上百个中间步骤——远超任何现有QA基准。为构建MoNaCo,我们开发了一种分解标注管道,以大规模引导和手动回答真实世界的耗时问题。评估基准中最先进的LLM在MoNaCo上最多只能达到61.2%的F1分数,受限于低召回率和幻觉。我们的结果凸显了LLM驱动的代理在处理复杂性和广度方面的局限性——MoNaCo为跟踪此类进展提供了有效资源。MoNaCo基准、代码库、提示和模型预测均可公开获取,地址为:https://tomerwolgithub.github.io/monaco
引用
@article{arxiv.2508.11133,
title = {MoNaCo: More Natural and Complex Questions for Reasoning Across Dozens of Documents},
author = {Tomer Wolfson and Harsh Trivedi and Mor Geva and Yoav Goldberg and Dan Roth and Tushar Khot and Ashish Sabharwal and Reut Tsarfaty},
journal= {arXiv preprint arXiv:2508.11133},
year = {2025}
}
备注
Accepted for publication in Transactions of the Association for Computational Linguistics (TACL), 2025. Authors pre-print