DETOUR:基于双智能体搜索与推理的交互式基准
计算与语言
2026-02-03 v1
摘要
在对话中回想信息时,人们往往经过多个回合后才能到达回想。然而,现有的基准受限于单回合设置,无法真实模拟 tip-of-the-tongue 搜索过程。为此,我们引入Dual-agent based Evaluation Through Obscure Under-specified Retrieval(DETOUR),一个包含1011个提示的双智能体评估基准。基准设计涉及Primary Agent和 Memory Agent两个角色。Primary Agent是评估对象,任务是通过查询 Memory Agent 来识别被回想的实体。我们的结果表明,当前最先进的模型仍在我们基准上仅能实现36%的准确率(在text、image、audio和video的所有模态上),这凸显了在不明确场景中提升能力的重要性。
引用
@article{arxiv.2602.00352,
title = {DETOUR: An Interactive Benchmark for Dual-Agent Search and Reasoning},
author = {Li Siyan and Darshan Deshpande and Anand Kannappan and Rebecca Qian},
journal= {arXiv preprint arXiv:2602.00352},
year = {2026}
}